Multi-Modal AI
Lectures: MW 3:30 - 5:00 PM
Room: Towne 337
Instructor: Jiayuan Mao
Generalist agents require multimodal understanding: observations can include
images, depth, IMU signals, motor state, audio, and instructions, while actions
may involve movement, manipulation, and communication. This course takes a
learning-based perspective on multimodal AI and asks how to formulate problems,
represent different modalities, align them, and reason across them.
The course studies representation and computation paradigms across visual,
language, audio, video, tactile, action, structured, and scientific domains.
Students will learn common principles such as contrastive learning,
cross-attention, generative objectives, multi-task learning, and tokenization.
Each class is organized as roughly one hour of lecture followed by about thirty
minutes of questions and discussion. Lectures are led by the instructor or guest
speakers and focus on topics such as 2D vision, multimodal architectures, and
large multimodal models.
Students should read the assigned papers before class, post discussion questions
online, and come prepared to connect each topic to their own research interests.
Students who want to present a paper should contact the instructor before class.
Projects may cover any topic related to multimodal AI, including topics connected
to a student's own research field. Teams may include up to two students, and
individual projects are welcome.
Master's theses and capstone projects are allowed with instructor approval. The
project may also be a subset of a paper for publication, as long as it is connected
to the course.
- Sun Sep 20: Initial proposal due.
- Sun Oct 11: Final proposal due.
- Sun Nov 1: Mid-project check-in through office hours.
- Wed Dec 2 and Mon Dec 7: Final project presentations.
| Component |
Weight |
Description |
| Participation |
25% |
Active engagement in lectures, discussions, and peer feedback. |
| Reading Responses |
20% |
Timely and thoughtful responses to assigned readings and discussion-question posts. |
| Project proposal |
10% |
Problem formulation, related work, proposed methodology, and evaluation plan. |
| Final project |
35% |
Research execution, experiments, analysis, and final report. |
| Final presentation |
10% |
Clear communication of motivation, method, results, and limitations. |
Dates follow the University of Pennsylvania Fall 2026 academic calendar.
The course meets Mondays and Wednesdays, 3:30 - 5:00 PM. Each weekly topic
normally spans two class meetings; holiday weeks are adjusted accordingly.
| Lec# |
Date |
Topic |
Reading |
| 1 |
Wed Aug 26 |
Course overview and multimodal AI formulation
Slides /
Recording
|
No reading |
| 2 |
Mon Aug 31 |
2D Vision
Slides /
Recording
|
Fully Convolutional Networks for Semantic Segmentation (2015)
Mask R-CNN (2017)
Unsupervised Learning of Depth and Ego-Motion from Video (2017)
An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale (2021)
Optional: Mask DINO: Towards A Unified Transformer-based Framework for Object Detection and Segmentation (2023)
|
| 3 |
Wed Sep 2 |
3D Vision
Slides /
Recording /
Recording (Zoom)
|
PointNet: Deep Learning on Point Sets for 3D Classification and Segmentation (2017)
VGGT: Visual Geometry Grounded Transformer (2025)
3D Dynamic Scene Graphs: Actionable Spatial Perception with Places, Objects, and Humans (2020)
Where2Act: From Pixels to Actions for Articulated 3D Objects (2021)
Optional: SayPlan: Grounding Large Language Models using 3D Scene Graphs for Scalable Robot Task Planning (2023)
|
|
Mon Sep 7 |
No class: Labor Day |
| 4 |
Wed Sep 9 |
Language Foundations
Slides /
Recording
|
Neural Machine Translation of Rare Words with Subword Units (2016)
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding (2019)
Language Models are Few-Shot Learners (2020)
Chain-of-Thought Prompting Elicits Reasoning in Large Language Models (2022)
Optional: SuperBPE: Space Travel for Language Models (2025)
|
| 5 |
Mon Sep 14 |
Vision-Language Learning
Slides /
Recording
|
VSE++: Improving Visual-Semantic Embeddings with Hard Negatives (2018)
Visual Genome: Connecting Language and Vision Using Crowdsourced Dense Image Annotations (2017)
Learning Transferable Visual Models From Natural Language Supervision (2021)
Winoground: Probing Vision and Language Models for Visio-Linguistic Compositionality (2022)
Optional: Align before Fuse: Vision and Language Representation Learning with Momentum Distillation (2021)
|
| 6 |
Wed Sep 16 |
Text-to-Image Generation
Slides /
Recording
|
Large Scale GAN Training for High Fidelity Natural Image Synthesis (2019)
Denoising Diffusion Probabilistic Models (2020)
High-Resolution Image Synthesis with Latent Diffusion Models (2022)
Classifier-Free Diffusion Guidance (2022)
Optional: Reduce, Reuse, Recycle: Compositional Generation with Energy-Based Diffusion Models and MCMC (2023)
|
|
Sun Sep 20 |
Project milestone: initial proposal due |
| 7 |
Mon Sep 21 |
Controllable and Structured Generation
Slides /
Recording
|
InstructPix2Pix: Learning to Follow Image Editing Instructions (2023)
Adding Conditional Control to Text-to-Image Diffusion Models (2023)
Compositional Image Decomposition with Diffusion Models (2024)
Scaling In-the-Wild Training for Diffusion-based Illumination Harmonization and Editing by Imposing Consistent Light Transport (2025)
Optional: Prompt-to-Prompt Image Editing with Cross Attention Control (2022)
|
| 8 |
Wed Sep 23 |
Speech and Audio Representation
Slides /
Recording
|
Deep Audio Priors Emerge From Harmonic Convolutional Networks (2020)
HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units (2021)
Hybrid Transformers for Music Source Separation (2022)
Anticipatory Music Transformer (2023)
Optional: Audio-Visual Neural Syntax Acquisition (2023)
|
| 9 |
Mon Sep 28 |
Speech and Audio Learning and Generation
Guest lecture:
Yunyun Wang (Adobe)
Online: instructor traveling
Slides /
Recording
|
No reading |
| 10 |
Wed Sep 30 |
Video Understanding
Online: instructor traveling
Slides /
Recording
|
TSM: Temporal Shift Module for Efficient Video Understanding (2019)
SlowFast Networks for Video Recognition (2019)
InternVideo2: Scaling Foundation Models for Multimodal Video Understanding (2024)
LongVideoBench: A Benchmark for Long-context Interleaved Video-Language Understanding (2024)
Optional: T*: Re-thinking Temporal Search for Long-Form Video Understanding (2025)
|
| 11 |
Mon Oct 5 |
Video Generation and Prediction |
VideoGPT: Video Generation using VQ-VAE and Transformers (2021)
CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer (2024)
Diffusion Forcing: Next-token Prediction Meets Full-Sequence Diffusion (2024)
History-Guided Video Diffusion (2025)
|
| 12 |
Wed Oct 7 |
Tokenization across Modalities |
From Pixels to Tokens: Byte-Pair Encoding on Quantized Visual Modalities (2024)
dMel: Speech Tokenization made Simple (2024)
UniTok: A Unified Tokenizer for Visual Generation and Understanding (2025)
An Image is Worth 32 Tokens for Reconstruction and Generation (2024)
Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning (2025)
|
|
Sun Oct 11 |
Project milestone: final proposal due |
| 13 |
Mon Oct 12 |
Discrete vs. Continuous Representations
Guest lecture:
Qihang Yu (Amazon FAR)
|
TBD |
| 14 |
Wed Oct 14 |
Connecting Modalities to LLMs |
TBD |
| 15 |
Mon Oct 19 |
Fusion Architectures and Unified Multimodal Models |
TBD |
| 16 |
Wed Oct 21 |
Scaling and Multimodal Data |
TBD |
| 17 |
Mon Oct 26 |
Pretraining, Post-Training, and Evaluation
Guest lecture:
Martin Ziqiao Ma
(Thinking Machines Lab)
|
TBD |
| 18 |
Wed Oct 28 |
Tactile Sensing and Learning
Guest lecture:
Binghao Huang (Columbia)
|
TBD |
|
Sun Nov 1 |
Project milestone: mid-project check-in through office hours |
| 19 |
Mon Nov 2 |
Action Representations |
TBD |
| 20 |
Wed Nov 4 |
Vision-Language-Action Models |
TBD |
| 21 |
Mon Nov 9 |
Charts, Tables, and Structured Domains
Online: instructor traveling
|
TBD |
| 22 |
Wed Nov 11 |
Biosignals, Biomedical Sensing, and Learning
Guest lecture:
Yingcheng Liu (MIT)
Online: instructor traveling
|
TBD |
| 23 |
Mon Nov 16 |
Compositionality and Neuro-Symbolic Models |
TBD |
| 24 |
Wed Nov 18 |
Planning and World Models |
TBD |
| 25 |
Mon Nov 23 |
Causal Discovery |
TBD |
|
Wed Nov 25 |
No regular Wednesday class: Friday class schedule |
| 26 |
Mon Nov 30 |
Agents, HCI, Privacy, Safety, and Alignment
Guest lecture:
Zora Wang (CMU)
|
TBD |
| 27 |
Wed Dec 2 |
Final Project Presentations |
TBD |
| 28 |
Mon Dec 7 |
Final Project Presentations; Last Day of Classes |
TBD |