Module 1: Introduction to Multimodal AI
- Lesson 1.1: The Shift from Unimodal to Multimodal Systems
- Learn how modern models bridge the gap between different data types.
- Lesson 1.2: Core Architectures and Foundations
- Understand the underlying tech behind CLIP, vision-language models (VLMs), and audio-speech transformers.
- Lesson 1.3: Embeddings Across Modalities
- Explore how to map text, images, and audio into a shared vector space