Course overview
This course aims to provide students with a comprehensive understanding of computer vision and multimodal machine learning techniques. Students will explore the theoretical foundations and practical applications of these technologies, gaining skills in designing, implementing, and evaluating models that integrate visual and multimodal data. The course will prepare students for advanced research or professional roles in fields that require sophisticated image and data analysis capabilities.
- Machine Learning for Computer Vision
- Multimodal Data Integration
- Advanced Topics in Computer Vision and Multimodal Machine Learning
Course learning outcomes
- Explain the fundamental concepts and applications of computer vision and multimodal machine learning.
- Preprocess images and extract meaningful features for analysis.
- Design, implement, and train convolutional neural networks (CNNs) for image classification and object detection tasks.
- Integrate and analyse multimodal data combining different types of data (e.g., text, image, and audio
- Apply advanced techniques such as generative models, transfer learning, and reinforcement learning to improve model performance.
- Develop and execute a comprehensive capstone project that demonstrates the ability to solve a complex problem using computer vision and multimodal machine learning techniques.