Get in Touch

Course Outline

Foundations of Multimodal LLMs in Vertex AI

  • Overview of multimodal capabilities within Vertex AI
  • Introduction to Gemini models and their supported modalities
  • Enterprise and research applications

Preparing the Development Environment

  • Configuring Vertex AI for multimodal workflows
  • Managing datasets across different modalities
  • Hands-on lab: setting up the environment and preparing datasets

Long Context Windows and Advanced Reasoning

  • Exploring the mechanics of long-context workflows
  • Applications in strategic planning and decision-making
  • Hands-on lab: executing long-context analysis tasks

Designing Cross-Modal Workflows

  • Synthesizing text, audio, and image analytics
  • Sequencing multimodal steps within a pipeline
  • Hands-on lab: architecting a multimodal pipeline

Configuring Gemini API Parameters

  • Setting up multimodal inputs and outputs
  • Optimizing inference speed and resource efficiency
  • Hands-on lab: fine-tuning Gemini API settings

Advanced Applications and System Integrations

  • Building interactive multimodal agents and assistants
  • Connecting external APIs and tools
  • Hands-on lab: developing a comprehensive multimodal application

Assessment and Continuous Improvement

  • Conducting performance tests for multimodal systems
  • Tracking metrics for accuracy, alignment, and data drift
  • Hands-on lab: evaluating the effectiveness of multimodal workflows

Key Takeaways and Future Directions

Requirements

  • Solid proficiency in Python programming
  • Hands-on experience in developing machine learning models
  • Working knowledge of multimodal data types (text, audio, and images)

Target Audience

  • AI researchers
  • Senior developers
  • Machine learning scientists
 14 Hours

Number of participants


Price per participant

Upcoming Courses

Related Categories