Login Sign Up

Introduction to Multimodal Prompting for Generative AI

Via LinkedIn Learning

LinkedIn Learning based on 144 ratings

Share

0

Overview

"Introduction to Multimodal Prompting for Generative AI" is an entry-level course designed for developers, AI enthusiasts, and content creators interested in leveraging multiple data types—such as text, images, and audio—in generative AI applications. This course explores how multimodal prompting enhances generative models' capabilities and allows for richer, more context-aware outputs. Participants will learn the foundational concepts behind multimodal models, including how they interpret and respond to diverse input formats. The course covers leading-edge tools and APIs that support multimodal prompts, including...

Syllabus

Introduction

  • GenAI with multimodal prompts

1. Multimodality

  • What is multimodality?
  • Visual modality
  • Textual and auditory modality

2. GPT-4

  • GPT-4 and 4o
  • Text to image in GPT-4
  • GPT-4 API with various input types
  • Challenge: Drawing to code
  • Solution: Drawing to code

3. Gemini

  • What is Gemini?
  • Images in Gemini
  • Gemini video inputs
  • Challenge: Video narration
  • Solution: Video narration

4. Auditory Modalities

  • Audio in generative AI
  • Prompt and audio
  • Generating music
  • Challenge: Soundtrack creation
  • Solution: Soundtrack creation

Conclusion

  • Next steps
Introduction to Multimodal Prompting for Generative AI
Go to Class

via LinkedIn Learning

43 minutes

Certificate Available

English

On-Demand

Instructor

Ronnie Sheer

Reviews

No reviews yet. Be the first to review!