Multi-Modal AI

Lectures: MW 3:30 - 5:00 PM
Room: Towne 337
Instructor: Jiayuan Mao

Course description

Generalist agents require multimodal understanding: observations can include images, depth, IMU signals, motor state, audio, and instructions, while actions may involve movement, manipulation, and communication. This course takes a learning-based perspective on multimodal AI and asks how to formulate problems, represent different modalities, align them, and reason across them.

The course studies representation and computation paradigms across visual, language, audio, video, tactile, action, structured, and scientific domains. Students will learn common principles such as contrastive learning, cross-attention, generative objectives, multi-task learning, and tokenization.

Course format

Each class is organized as roughly one hour of lecture followed by about thirty minutes of questions and discussion. Lectures are led by the instructor or guest speakers and focus on topics such as 2D vision, multimodal architectures, and large multimodal models.

Students should read the assigned papers before class, post discussion questions online, and come prepared to connect each topic to their own research interests. Students who want to present a paper should contact the instructor before class.

Final project

Projects may cover any topic related to multimodal AI, including topics connected to a student's own research field. Teams may include up to two students, and individual projects are welcome.

Master's theses and capstone projects are allowed with instructor approval. The project may also be a subset of a paper for publication, as long as it is connected to the course.

Assessment
Component Weight Description
Participation 25% Active engagement in lectures, discussions, and peer feedback.
Reading Responses 20% Timely and thoughtful responses to assigned readings and discussion-question posts.
Project proposal 10% Problem formulation, related work, proposed methodology, and evaluation plan.
Final project 35% Research execution, experiments, analysis, and final report.
Final presentation 10% Clear communication of motivation, method, results, and limitations.
Schedule

Dates follow the University of Pennsylvania Fall 2026 academic calendar. The course meets Mondays and Wednesdays, 3:30 - 5:00 PM. Each weekly topic normally spans two class meetings; holiday weeks are adjusted accordingly.

Lec# Date Topic Reading
1 Wed Aug 26 Course overview and multimodal AI formulation
Slides / Recording
No reading
2 Mon Aug 31 2D Vision
Slides / Recording
Fully Convolutional Networks for Semantic Segmentation (2015)
Mask R-CNN (2017)
Unsupervised Learning of Depth and Ego-Motion from Video (2017)
An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale (2021)
Optional: Mask DINO: Towards A Unified Transformer-based Framework for Object Detection and Segmentation (2023)
3 Wed Sep 2 3D Vision PointNet: Deep Learning on Point Sets for 3D Classification and Segmentation (2017)
VGGT: Visual Geometry Grounded Transformer (2025)
3D Dynamic Scene Graphs: Actionable Spatial Perception with Places, Objects, and Humans (2020)
Where2Act: From Pixels to Actions for Articulated 3D Objects (2021)
Optional: SayPlan: Grounding Large Language Models using 3D Scene Graphs for Scalable Robot Task Planning (2023)
Mon Sep 7 No class: Labor Day
4 Wed Sep 9 Language Foundations Neural Machine Translation of Rare Words with Subword Units (2016)
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding (2019)
Language Models are Few-Shot Learners (2020)
Chain-of-Thought Prompting Elicits Reasoning in Large Language Models (2022)
Optional: SuperBPE: Space Travel for Language Models (2025)
5 Mon Sep 14 Vision-Language Learning TBD
6 Wed Sep 16 Text-to-Image Generation TBD
Sun Sep 20 Project milestone: initial proposal due
7 Mon Sep 21 Controllable and Structured Generation TBD
8 Wed Sep 23 Speech and Audio Representation TBD
9 Mon Sep 28 Speech and Audio Learning and Generation
Online: instructor traveling
TBD
10 Wed Sep 30 Video Understanding
Online: instructor traveling
TBD
11 Mon Oct 5 Video Generation and Prediction TBD
12 Wed Oct 7 Tokenization across Modalities TBD
Sun Oct 11 Project milestone: final proposal due
13 Mon Oct 12 Discrete vs. Continuous Representations
Guest lecture: Qihang Yu (Amazon FAR)
TBD
14 Wed Oct 14 Connecting Modalities to LLMs TBD
15 Mon Oct 19 Fusion Architectures and Unified Multimodal Models TBD
16 Wed Oct 21 Scaling and Multimodal Data TBD
17 Mon Oct 26 Pretraining, Post-Training, and Evaluation
Guest lecture: Martin Ziqiao Ma (Thinking Machines Lab)
TBD
18 Wed Oct 28 Tactile Sensing and Learning
Guest lecture: Binghao Huang (Columbia)
TBD
Sun Nov 1 Project milestone: mid-project check-in through office hours
19 Mon Nov 2 Action Representations TBD
20 Wed Nov 4 Vision-Language-Action Models TBD
21 Mon Nov 9 Charts, Tables, and Structured Domains
Online: instructor traveling
TBD
22 Wed Nov 11 Biosignals, Biomedical Sensing, and Learning
Guest lecture: Yingcheng Liu (MIT)
Online: instructor traveling
TBD
23 Mon Nov 16 Compositionality and Neuro-Symbolic Models TBD
24 Wed Nov 18 Planning and World Models TBD
25 Mon Nov 23 Causal Discovery TBD
Wed Nov 25 No regular Wednesday class: Friday class schedule
26 Mon Nov 30 Agents, HCI, Privacy, Safety, and Alignment
Guest lecture: Zora Wang (CMU)
TBD
27 Wed Dec 2 Final Project Presentations TBD
28 Mon Dec 7 Final Project Presentations; Last Day of Classes TBD