Login Sign Up

Multimodal Generative AI: Vision, Speech, and Assistants

Codio via Coursera

Coursera based on 0 ratings

Share

0

Overview

The "Multimodal Generative AI: Vision, Speech, and Assistants" course introduces learners to the exciting frontier of AI systems that understand and generate content across multiple modalities—such as text, images, audio, and video. This course is ideal for those interested in how cutting-edge AI powers tools like virtual assistants, image captioning systems, voice generators, and AI chatbots that integrate vision and speech. You will explore the foundations of multimodal AI, including how it combines data from different sources and uses models like...

Syllabus

  • Image to text
    • Welcome to Week 1 of the course. These assignments cover vision and image-to-text capabilities. You'll learn how to analyze and interpret images using AI. The module ends with graded summative assessments.
  • Text to Speech
    • Welcome to Week 2 of the course. This week focuses on understanding the fundamentals of text-to-speech (TTS). These assignments cover generating spoken audio in different voices. The module ends with graded summative assessments.
  • Speech to Text
    • Welcome to Week 3 of the course. You'll understand the basics of Whisper and interact with ChatGPT to enhance and optimize the Whisper API. The module ends with graded summative assessments.
  • Assistants
    • Welcome to Week 4 of the course. These assignments cover understanding the basics of the Assistants API, including their purpose, primary components, and available tools like Code Interpreter, File Search, and Function Calling. The module ends with graded summative assessments.
Multimodal Generative AI: Vision, Speech, and Assistants
Go to Class

Codio via Coursera

14 hours 25 minutes

Paid Certificate Available

English

On-Demand

Beginner

Instructor

Kevin Noelsaint

Reviews

No reviews yet. Be the first to review!