Lectures: MW 3:30 - 5:00 PM
Room: Towne 337
Instructor: Jiayuan Mao
Generalist agents require multimodal understanding: observations can include images, depth, IMU signals, motor state, audio, and instructions, while actions may involve movement, manipulation, and communication. This course takes a learning-based perspective on multimodal AI and asks how to formulate problems, represent different modalities, align them, and reason across them.
The course studies representation and computation paradigms across visual, language, audio, video, tactile, action, structured, and scientific domains. Students will learn common principles such as contrastive learning, cross-attention, generative objectives, multi-task learning, and tokenization.
Each class is organized as roughly one hour of lecture followed by about thirty minutes of questions and discussion. Lectures are led by the instructor or guest speakers and focus on topics such as 2D vision, multimodal architectures, and large multimodal models.
Students should read the assigned papers before class, post discussion questions online, and come prepared to connect each topic to their own research interests. Students who want to present a paper should contact the instructor before class.
Projects may cover any topic related to multimodal AI, including topics connected to a student's own research field. Teams may include up to two students, and individual projects are welcome.
Master's theses and capstone projects are allowed with instructor approval. The project may also be a subset of a paper for publication, as long as it is connected to the course.
| Component | Weight | Description |
|---|---|---|
| Participation | 25% | Active engagement in lectures, discussions, and peer feedback. |
| Reading Responses | 20% | Timely and thoughtful responses to assigned readings and discussion-question posts. |
| Project proposal | 10% | Problem formulation, related work, proposed methodology, and evaluation plan. |
| Final project | 35% | Research execution, experiments, analysis, and final report. |
| Final presentation | 10% | Clear communication of motivation, method, results, and limitations. |
| Week | Theme | Topics |
|---|---|---|
| 1 | Foundations | Overview of multimodal AI; representation, alignment, and fusion. |
| 2 | 2D/3D Vision and Perception | Visual representations; 3D and spatial perception. |
| 3 | Vision and Text | Language foundations; vision-language learning. |
| 4 | Text-to-Image | Generative vision; controllable and structured generation. |
| 5 | Audio | Speech/audio representation; audio-language and generation. |
| 6 | Video | Video understanding; video generation and prediction. |
| 7 | Tokenization | Tokenization across modalities; discrete vs. continuous representations. |
| 8 | Multimodal Architectures | Connecting modalities to LLMs; fusion architectures; unified multimodal models. |
| 9 | Large Multimodal Models | Scaling, multimodal data, pretraining, post-training, and evaluation. |
| 10 | Structured & Scientific Domains | Charts, tables, and scientific modalities. |
| 11 | Tactile and Action | Tactile/sensor learning; action representations; vision-language-action models. |
| 12 | Structure & Reasoning | Grounding, planning, causal discovery, and world models. |
| 13 | Advanced Topics | Agents, HCI, privacy, safety, and alignment. |