Thank you for sending your enquiry! One of our team members will contact you shortly.
Thank you for sending your booking! One of our team members will contact you shortly.
Course Outline
Foundations of Multimodal LLMs in Vertex AI
- Overview of multimodal capabilities within Vertex AI
- Introduction to Gemini models and their supported modalities
- Enterprise and research applications
Preparing the Development Environment
- Configuring Vertex AI for multimodal workflows
- Managing datasets across different modalities
- Hands-on lab: setting up the environment and preparing datasets
Long Context Windows and Advanced Reasoning
- Exploring the mechanics of long-context workflows
- Applications in strategic planning and decision-making
- Hands-on lab: executing long-context analysis tasks
Designing Cross-Modal Workflows
- Synthesizing text, audio, and image analytics
- Sequencing multimodal steps within a pipeline
- Hands-on lab: architecting a multimodal pipeline
Configuring Gemini API Parameters
- Setting up multimodal inputs and outputs
- Optimizing inference speed and resource efficiency
- Hands-on lab: fine-tuning Gemini API settings
Advanced Applications and System Integrations
- Building interactive multimodal agents and assistants
- Connecting external APIs and tools
- Hands-on lab: developing a comprehensive multimodal application
Assessment and Continuous Improvement
- Conducting performance tests for multimodal systems
- Tracking metrics for accuracy, alignment, and data drift
- Hands-on lab: evaluating the effectiveness of multimodal workflows
Key Takeaways and Future Directions
Requirements
- Solid proficiency in Python programming
- Hands-on experience in developing machine learning models
- Working knowledge of multimodal data types (text, audio, and images)
Target Audience
- AI researchers
- Senior developers
- Machine learning scientists
14 Hours