Lectures: MW 3:30 - 5:00 PM
Room: Towne 337
Instructor: Jiayuan Mao
Generalist agents require multimodal understanding: observations can include images, depth, IMU signals, motor state, audio, and instructions, while actions may involve movement, manipulation, and communication. This course takes a learning-based perspective on multimodal AI and asks how to formulate problems, represent different modalities, align them, and reason across them.
The course studies representation and computation paradigms across visual, language, audio, video, tactile, action, structured, and scientific domains. Students will learn common principles such as contrastive learning, cross-attention, generative objectives, multi-task learning, and tokenization.
Each class is organized as roughly one hour of lecture followed by about thirty minutes of questions and discussion. Lectures are led by the instructor or guest speakers and focus on topics such as 2D vision, multimodal architectures, and large multimodal models.
Students should read the assigned papers before class, post discussion questions online, and come prepared to connect each topic to their own research interests. Students who want to present a paper should contact the instructor before class.
Projects may cover any topic related to multimodal AI, including topics connected to a student's own research field. Teams may include up to two students, and individual projects are welcome.
Master's theses and capstone projects are allowed with instructor approval. The project may also be a subset of a paper for publication, as long as it is connected to the course.
| Component | Weight | Description |
|---|---|---|
| Participation | 25% | Active engagement in lectures, discussions, and peer feedback. |
| Reading Responses | 20% | Timely and thoughtful responses to assigned readings and discussion-question posts. |
| Project proposal | 10% | Problem formulation, related work, proposed methodology, and evaluation plan. |
| Final project | 35% | Research execution, experiments, analysis, and final report. |
| Final presentation | 10% | Clear communication of motivation, method, results, and limitations. |
Dates follow the University of Pennsylvania Fall 2026 academic calendar. The course meets Mondays and Wednesdays, 3:30 - 5:00 PM. Each weekly topic normally spans two class meetings; holiday weeks are adjusted accordingly.
| Lec# | Date | Topic | Reading |
|---|---|---|---|
| 1 | Wed Aug 26 |
Course overview and multimodal AI formulation Slides / Recording |
No reading |
| 2 | Mon Aug 31 |
2D Vision Slides / Recording |
Fully Convolutional Networks for Semantic Segmentation (2015) Mask R-CNN (2017) Unsupervised Learning of Depth and Ego-Motion from Video (2017) An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale (2021) Optional: Mask DINO: Towards A Unified Transformer-based Framework for Object Detection and Segmentation (2023) |
| 3 | Wed Sep 2 | 3D Vision |
PointNet: Deep Learning on Point Sets for 3D Classification and Segmentation (2017) VGGT: Visual Geometry Grounded Transformer (2025) 3D Dynamic Scene Graphs: Actionable Spatial Perception with Places, Objects, and Humans (2020) Where2Act: From Pixels to Actions for Articulated 3D Objects (2021) Optional: SayPlan: Grounding Large Language Models using 3D Scene Graphs for Scalable Robot Task Planning (2023) |
| Mon Sep 7 | No class: Labor Day | ||
| 4 | Wed Sep 9 | Language Foundations |
Neural Machine Translation of Rare Words with Subword Units (2016) BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding (2019) Language Models are Few-Shot Learners (2020) Chain-of-Thought Prompting Elicits Reasoning in Large Language Models (2022) Optional: SuperBPE: Space Travel for Language Models (2025) |
| 5 | Mon Sep 14 | Vision-Language Learning | TBD |
| 6 | Wed Sep 16 | Text-to-Image Generation | TBD |
| Sun Sep 20 | Project milestone: initial proposal due | ||
| 7 | Mon Sep 21 | Controllable and Structured Generation | TBD |
| 8 | Wed Sep 23 | Speech and Audio Representation | TBD |
| 9 | Mon Sep 28 |
Speech and Audio Learning and Generation Online: instructor traveling |
TBD |
| 10 | Wed Sep 30 |
Video Understanding Online: instructor traveling |
TBD |
| 11 | Mon Oct 5 | Video Generation and Prediction | TBD |
| 12 | Wed Oct 7 | Tokenization across Modalities | TBD |
| Sun Oct 11 | Project milestone: final proposal due | ||
| 13 | Mon Oct 12 |
Discrete vs. Continuous Representations Guest lecture: Qihang Yu (Amazon FAR) |
TBD |
| 14 | Wed Oct 14 | Connecting Modalities to LLMs | TBD |
| 15 | Mon Oct 19 | Fusion Architectures and Unified Multimodal Models | TBD |
| 16 | Wed Oct 21 | Scaling and Multimodal Data | TBD |
| 17 | Mon Oct 26 |
Pretraining, Post-Training, and Evaluation Guest lecture: Martin Ziqiao Ma (Thinking Machines Lab) |
TBD |
| 18 | Wed Oct 28 |
Tactile Sensing and Learning Guest lecture: Binghao Huang (Columbia) |
TBD |
| Sun Nov 1 | Project milestone: mid-project check-in through office hours | ||
| 19 | Mon Nov 2 | Action Representations | TBD |
| 20 | Wed Nov 4 | Vision-Language-Action Models | TBD |
| 21 | Mon Nov 9 |
Charts, Tables, and Structured Domains Online: instructor traveling |
TBD |
| 22 | Wed Nov 11 |
Biosignals, Biomedical Sensing, and Learning Guest lecture: Yingcheng Liu (MIT) Online: instructor traveling |
TBD |
| 23 | Mon Nov 16 | Compositionality and Neuro-Symbolic Models | TBD |
| 24 | Wed Nov 18 | Planning and World Models | TBD |
| 25 | Mon Nov 23 | Causal Discovery | TBD |
| Wed Nov 25 | No regular Wednesday class: Friday class schedule | ||
| 26 | Mon Nov 30 |
Agents, HCI, Privacy, Safety, and Alignment Guest lecture: Zora Wang (CMU) |
TBD |
| 27 | Wed Dec 2 | Final Project Presentations | TBD |
| 28 | Mon Dec 7 | Final Project Presentations; Last Day of Classes | TBD |