Module 2: Computer Vision and Text Integration
- Lesson 2.1: Image Captioning and Visual Question Answering (VQA)
- Build systems that can describe images and answer specific questions about visual data.
- Lesson 2.2: Cross-Modal Retrieval
- Develop search engines that allow users to search for images using natural language text queries.
- Lesson 2.3: Zero-Shot Image Classification
- Deploy models to classify images into custom categories without needing traditional retraining.