Login Sign Up

Build an Image Captioning Tool for Visually Impaired Users with Gemini

Via LinkedIn Learning

LinkedIn Learning based on 7 ratings

Share

0

Overview

“Build an Image Captioning Tool for Visually Impaired Users with Gemini” is a hands-on course designed to guide developers in creating accessible AI applications using Google’s Gemini multimodal models. The course focuses on leveraging Gemini’s advanced image recognition and natural language generation capabilities to build an intuitive image captioning tool that aids visually impaired users in understanding visual content. Participants will learn how to process images through Gemini’s APIs, generate accurate and descriptive captions, and optimize output for clarity and...

Syllabus

Introduction

  • Image captioning with AI
  • What you should know
  • Who this course is for

1. Setting Up Access to Gemini API

  • Understanding Gemini models
  • Gemini pricing
  • Signing up for an Google AI Studio account
  • Getting your API key

2. Building the Interface

  • Cloning the seed project
  • Project code walkthrough
  • Adding the image upload functionality
  • Adding the prompt functionality
  • Writing the caption display

3. Building the Backend: Connecting to Gemini

  • Building out the Express.js API
  • Configuring the Generative AI SDK
  • Adding routes
  • Setting up file upload functionality
  • Writing the prompt request and response

4. Bringing It All Together

  • Connecting the frontend to the API
  • Adding a progress indicator
  • Using the Web Speech API to read captions

Conclusion

  • Next steps
Build an Image Captioning Tool for Visually Impaired Users with Gemini
Go to Class

via LinkedIn Learning

1 hour 8 minutes

Certificate Available

English

On-Demand

Instructor

Fikayo Adepoju

Reviews

No reviews yet. Be the first to review!