
Vision AI and Advanced OpenAI
coursera · Datos e IA · en
Impartido por Mumshad Mannambeth
Precio no disponible en esta plataforma
Entra en tu cuenta para guardar este curso y volver a él cuando quieras.
Comparar este cursoEnlace de afiliado: podemos cobrar comisión, sin coste extra para ti. Más información
Descripción
Once you can build with text, the next step is multimodal AI and enterprise-scale capabilities. This course advances your OpenAI skills into image generation, advanced production API features, and the foundational AWS architecture knowledge you need to design serious generative AI systems. You'll start with OpenAI's vision capabilities: how DALL-E generates images from natural language prompts, the evolution from DALL-E 1 to DALL-E 3, and how CLIP connects visual and language understanding. You'll build a working image generator and an image captioning pipeline, and examine the ethical challenges and future trends shaping AI vision technologies. Next, you'll implement OpenAI's most powerful production features — function calling, structured outputs, batch processing, and content moderation — the capabilities that separate prototype applications from scalable, enterprise-grade systems. The course closes with a critical knowledge bridge into AWS: tokens, embeddings, chunking, context windows, the foundation model lifecycle, AWS GenAI infrastructure design, and cost optimisation principles — preparing you for AWS-native deployment. Designed for learners who have OpenAI API experience. Basic programming knowledge is recommended.