Multi-Modal AI

Lectures: MW 3:30 - 5:00 PM
Room: Towne 337
Instructor: Jiayuan Mao

Course description

Generalist agents require multimodal understanding: observations can include images, depth, IMU signals, motor state, audio, and instructions, while actions may involve movement, manipulation, and communication. This course takes a learning-based perspective on multimodal AI and asks how to formulate problems, represent different modalities, align them, and reason across them.

The course studies representation and computation paradigms across visual, language, audio, video, tactile, action, structured, and scientific domains. Students will learn common principles such as contrastive learning, cross-attention, generative objectives, multi-task learning, and tokenization.

Course format

Each class is organized as roughly one hour of lecture followed by about thirty minutes of questions and discussion. Lectures are led by the instructor or guest speakers and focus on topics such as 2D vision, multimodal architectures, and large multimodal models.

Students should read the assigned papers before class, post discussion questions online, and come prepared to connect each topic to their own research interests. Students who want to present a paper should contact the instructor before class.

Final project

Projects may cover any topic related to multimodal AI, including topics connected to a student's own research field. Teams may include up to two students, and individual projects are welcome.

Master's theses and capstone projects are allowed with instructor approval. The project may also be a subset of a paper for publication, as long as it is connected to the course.

Assessment
Component Weight Description
Participation 25% Active engagement in lectures, discussions, and peer feedback.
Reading Responses 20% Timely and thoughtful responses to assigned readings and discussion-question posts.
Project proposal 10% Problem formulation, related work, proposed methodology, and evaluation plan.
Final project 35% Research execution, experiments, analysis, and final report.
Final presentation 10% Clear communication of motivation, method, results, and limitations.
Schedule (tentative)
Week Theme Topics
1 Foundations Overview of multimodal AI; representation, alignment, and fusion.
2 2D/3D Vision and Perception Visual representations; 3D and spatial perception.
3 Vision and Text Language foundations; vision-language learning.
4 Text-to-Image Generative vision; controllable and structured generation.
5 Audio Speech/audio representation; audio-language and generation.
6 Video Video understanding; video generation and prediction.
7 Tokenization Tokenization across modalities; discrete vs. continuous representations.
8 Multimodal Architectures Connecting modalities to LLMs; fusion architectures; unified multimodal models.
9 Large Multimodal Models Scaling, multimodal data, pretraining, post-training, and evaluation.
10 Structured & Scientific Domains Charts, tables, and scientific modalities.
11 Tactile and Action Tactile/sensor learning; action representations; vision-language-action models.
12 Structure & Reasoning Grounding, planning, causal discovery, and world models.
13 Advanced Topics Agents, HCI, privacy, safety, and alignment.