pliang279/awesome-multimodal-ml
Reading list for research topics in multimodal machine learning
About pliang279/awesome-multimodal-ml
pliang279/awesome-multimodal-ml is an open-source project on GitHub, mainly written in several languages. Reading list for research topics in multimodal machine learning It currently holds 6,931 stars and 0 forks with 0 open issues, and was last pushed on an unknown date (repository created unknown).
Project Overview
AI Homed tracks it on the AI Image Projects board and on the AI AI Image Projects list.
GitHub Repository Details
README
Awesome Multimodal Machine Learning
By Paul Liang (pliang@cs.cmu.edu), Machine Learning Department and Language Technologies Institute, CMU, with help from members of the MultiComp Lab at LTI, CMU. If there are any areas, papers, and datasets I missed, please let me know!
Course content + workshops
Check out our comprehsensive tutorial paper Foundations and Recent Trends in Multimodal Machine Learning: Principles, Challenges, and Open Questions.
Tutorials on Multimodal Machine Learning at CVPR 2022 and NAACL 2022, slides and videos here.
New course 11-877 Advanced Topics in Multimodal Machine Learning Spring 2022 @ CMU. It will primarily be reading and discussion-based. We plan to post discussion probes, relevant papers, and summarized discussion highlights every week on the website.
Public course content and lecture videos from 11-777 Multimodal Machine Learning, Fall 2020 @ CMU.
Table of Contents
- Survey Papers
- Core Areas
- Multimodal Representations
- Multimodal Fusion
- Multimodal Alignment
- Multimodal Pretraining
- Multimodal Translation
- Crossmodal Retrieval
- Multimodal Co-learning
- Missing or Imperfect Modalities
- Analysis of Multimodal Models
- Knowledge Graphs and Knowledge Bases
- Intepretable Learning
- Generative Learning
- Semi-supervised Learning
- Self-supervised Learning
- Language Models
- Adversarial Attacks
- Few-Shot Learning
- Bias and Fairness
- Human in the Loop Learning
- Architectures
- Multimodal Transformers
- Multimodal Memory
- Applications and Datasets
- Language and Visual QA
- Language Grounding in Vision
- Language Grouding in Navigation
- Multimodal Machine Translation
- Multi-agent Communication
- Commonsense Reasoning
- Multimodal Reinforcement Learning
- Multimodal Dialog
- Language and Audio
- Audio and Visual
- Visual, IMU and Wireless
- Media Description
- Video Generation from Text
- Affect Recognition and Multimodal Language
- Healthcare
- Robotics
- Autonomous Driving
- Finance
- Human AI Interaction
- Workshops
- Tutorials
- Courses
Research Papers
Survey Papers
Foundations and Trends in Multimodal Machine Learning: Principles, Challenges, and Open Questions, arxiv 2023
Multimodal Learning with Transformers: A Survey, TPAMI 2023
Trends in Integration of Vision and Language Research: A Survey of Tasks, Datasets, and Methods, JAIR 2021
Experience Grounds Language, EMNLP 2020
A Survey of Reinforcement Learning Informed by Natural Language, IJCAI 2019
Multimodal Machine Learning: A Survey and Taxonomy, TPAMI 2019
Multimodal Intelligence: Representation Learning, Information Fusion, and Applications, arXiv 2019
Deep Multimodal Representation Learning: A Survey, arXiv 2019
Guest Editorial: Image and Language Understanding, IJCV 2017
Representation Learning: A Review and New Perspectives, TPAMI 2013
A Survey of Socially Interactive Robots, 2003
Core Areas
Multimodal Representations
Identifiability Results for Multimodal Contrastive Learning, ICLR 2023 [[code]](https://github.com/imantdaunhawer/multimodal-contrastive-learning)
Unpaired Vision-Language Pre-training via Cross-Modal CutMix, ICML 2022.
Balanced Multimodal Learning via On-the-fly Gradient Modulation, CVPR 2022
Unsupervised Voice-Face Representation Learning by Cross-Modal Prototype Contrast, IJCAI 2021 [[code]](https://github.com/Cocoxili/CMPC)
Towards a Unified Foundation Model: Jointly Pre-Training Transformers on Unpaired Images and Text, arXiv 2021
FLAVA: A Foundational Language And Vision Alignment Model, arXiv 2021
Transformer is All You Need: Multimodal Multitask Learning with a Unified Transformer, arXiv 2021
MultiBench: Multiscale Benchmarks for Multimodal Representation Learning, NeurIPS 2021 [[code]](https://github.com/pliang279/MultiBench)
Perceiver: General Perception with Iterative Attention, ICML 2021 [[code]](https://github.com/deepmind/deepmind-research/tree/master/perceiver)
Learning Transferable Visual Models From Natural Language Supervision, arXiv 2021 [[blog]](blog) [[code]](https://github.com/OpenAI/CLIP)
VinVL: Revisiting Visual Representations in Vision-Language Models, arXiv 2021 [[blog]](https://www.microsoft.com/en-us/research/blog/vinvl-advancing-the-state-of-the-art-for-vision-language-models/?OCID=msr_blog_VinVL_fb) [[code]](https://github.com/pzzhang/VinVL)
Learning Transferable Visual Models From Natural Language Supervision, arXiv 2020 [[blog]](https://openai.com/blog/clip/) [[code]](https://github.com/openai/CLIP)
12-in-1: Multi-Task Vision and Language Representation Learning, CVPR 2020 [[code]](https://github.com/facebookresearch/vilbert-multi-task)
Watching the World Go By: Representation Learning from Unlabeled Videos, arXiv 2020
Learning Video Representations using Contrastive Bidirectional Transformer, arXiv 2019
Visual Concept-Metaconcept Learning, NeurIPS 2019 [[code]](http://vcml.csail.mit.edu/)
OmniNet: A Unified Architecture for Multi-modal Multi-task Learning, arXiv 2019 [[code]](https://github.com/subho406/OmniNet)
Learning Representations by Maximizing Mutual Information Across Views, arXiv 2019 [[code]](https://github.com/Philip-Bachman/amdim-public)
ViCo: Word Embeddings from Visual Co-occurrences, ICCV 2019 [[code]](https://github.com/BigRedT/vico)
Multi-Task Learning of Hierarchical Vision-Language Representation, CVPR 2019
Learning Factorized Multimodal Representations, ICLR 2019 [[code]](https://github.com/pliang279/factorized/)
Do Neural Network Cross-Modal Mappings Really Bridge Modalities?, ACL 2018
Learning Robust Visual-Semantic Embeddings, ICCV 2017
Deep Multimodal Representation Learning from Temporal Data, CVPR 2017
Combining Language and Vision with a Multimodal Skip-gram Model, NAACL 2015
Deep Fragment Embeddings for Bidirectional Image Sentence Mapping, NIPS 2014
Multimodal Learning with Deep Boltzmann Machines, JMLR 2014
Learning Grounded Meaning Representations with Autoencoders, ACL 2014
DeViSE: A Deep Visual-Semantic Embedding Model, NeurIPS 2013
Multimodal Deep Learning, ICML 2011
Multimodal Fusion
Robust Contrastive Learning against Noisy Views, arXiv 2022
Cooperative Learning for Multi-view Analysis, arXiv 2022
What Makes Multi-modal Learning Better than Single (Provably), NeurIPS 2021
Efficient Multi-Modal Fusion with Diversity Analysis, ACMMM 2021
Attention Bottlenecks for Multimodal Fusion, NeurIPS 2021
VMLoc: Variational Fusion For Learning-Based Multimodal Camera Localization, AAAI 2021
Trusted Multi-View Classification, ICLR 2021 [[code]](https://github.com/hanmenghan/TMC)
Deep-HOSeq: Deep Higher-Order Sequence Fusion for Multimodal Sentiment Analysis, ICDM 2020
Removing Bias in Multi-modal Classifiers: Regularization by Maximizing Functional Entropies, NeurIPS 2020 [[code]](https://github.com/itaigat/removing-bias-in-multi-modal-classifiers)
Deep Multimodal Fusion by Channel Exchanging, NeurIPS 2020 [[code]](https://github.com/yikaiw/CEN)
What Makes Training Multi-Modal Classification Networks Hard?, CVPR 2020
Dynamic Fusion for Multimodal Data, arXiv 2019
DeepCU: Integrating Both Common and Unique Latent Information for Multimodal Sentiment Analysis, IJCAI 2019 [[code]](https://github.com/sverma88/DeepCU-IJCAI19)
Deep Multimodal Multilinear Fusion with High-order Polynomial Pooling, NeurIPS 2019
XFlow: Cross-modal Deep Neural Networks for Audiovisual Classification, IEEE TNNLS 2019 [[code]](https://github.com/catalina17/XFlow)
MFAS: Multimodal Fusion Architecture Search, CVPR 2019
The Neuro-Symbolic Concept Learner: Interpreting Scenes, Words, and Sentences From Natural Supervision, ICLR 2019 [[code]](http://nscl.csail.mit.edu/)
Unifying and merging well-trained deep neural networks for inference stage, IJCAI 2018 [[code]](https://github.com/ivclab/NeuralMerger)
Efficient Low-rank Multimodal Fusion with Modality-Specific Factors, ACL 2018 [[code]](https://github.com/Justin1904/Low-rank-Multimodal-Fusion)
Memory Fusion Network for Multi-view Sequential Learning, AAAI 2018 [[code]](https://github.com/pliang279/MFN)
Tensor Fusion Network for Multimodal Sentiment Analysis, EMNLP 2017 [[code]](https://github.com/A2Zadeh/TensorFusionNetwork)
A co-regularized approach to semi-supervised learning with multiple views, ICML 2005
Multimodal Alignment
Reconsidering Representation Alignment for Multi-view Clustering, CVPR 2021 [[code]](https://github.com/DanielTrosten/mvc)
CoMIR: Contrastive Multimodal Image Representation for Registration, NeurIPS 2020 [[code]](https://github.com/MIDA-group/CoMIR)
Multimodal Transformer for Unaligned Multimodal Language Sequences, ACL 2019 [[code]](https://github.com/yaohungt/Multimodal-Transformer)
Temporal Cycle-Consistency Learning, CVPR 2019 [[code]](https://github.com/google-research/google-research/tree/master/tcc)
See, Hear, and Read: Deep Aligned Representations, arXiv 2017
On Deep Multi-View Representation Learning, ICML 2015
Unsupervised Alignment of Natural Language Instructions with Video Segments, AAAI 2014
Multimodal Alignment of Videos, MM 2014
Deep Canonical Correlation Analysis, ICML 2013 [[code]](https://github.com/VahidooX/DeepCCA)
Multimodal Pretraining
Align before Fuse: Vision and Language Representation Learning with Momentum Distillation, NeurIPS 2021 Spotlight [[code]](https://github.com/salesforce/ALBEF)Less is More: ClipBERT for Video-and-Language Learning via Sparse Sampling, CVPR 2021 [[code]](https://github.com/jayleicn/ClipBERT)
Transformer is All You Need: Multimodal Multitask Learning with a Unified Transformer, arXiv 2021
Large-Scale Adversarial Training for Vision-and-Language Representation Learning, NeurIPS 2020 [[code]](https://github.com/zhegan27/VILLA)
Vokenization: Improving Language Understanding with Contextualized, Visual-Grounded Supervision, EMNLP 2020 [[code]](https://github.com/airsplay/vokenization)
Integrating Multimodal Information in Large Pretrained Transformers, ACL 2020
VL-BERT: Pre-training of Generic Visual-Linguistic Representations, arXiv 2019 [[code]](https://github.com/jackroos/VL-BERT)
VisualBERT: A Simple and Performant Baseline for Vision and Language, arXiv 2019 [[code]](https://github.com/uclanlp/visualbert)
ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks, NeurIPS 2019 [[code]](https://github.com/jiasenlu/vilbert_beta)
Unicoder-VL: A Universal Encoder for Vision and Language by Cross-modal Pre-training, arXiv 2019
LXMERT: Learning Cross-Modality Encoder Representations from Transformers, EMNLP 2019 [[code]](https://github.com/airsplay/lxmert)
VideoBERT: A Joint Model for Video and Language Representation Learning, ICCV 2019
Multimodal Translation
Zero-Shot Text-to-Image Generation, ICML 2021 [[code]](https://github.com/openai/DALL-E)
Translate-to-Recognize Networks for RGB-D Scene Recognition, CVPR 2019 [[code]](https://github.com/ownstyledu/Translate-to-Recognize-Networks)
Language2Pose: Natural Language Grounded Pose Forecasting, 3DV 2019 [[code]](http://chahuja.com/language2pose/)
Reconstructing Faces from Voices, NeurIPS 2019 [[code]](https://github.com/cmu-mlsp/reconstructing_faces_from_voices)
Speech2Face: Learning the Face Behind a Voice, CVPR 2019 [[code]](https://speech2face.github.io/)
Found in Translation: Learning Robust Joint Representations by Cyclic Translations Between Modalities, AAAI 2019 [[code]](https://github.com/hainow/MCTN)
Natural TTS Synthesis by Conditioning Wavenet on Mel Spectrogram Predictions, ICASSP 2018 [[code]](https://github.com/NVIDIA/tacotron2)
Crossmodal Retrieval
Learning with Noisy Correspondence for Cross-modal Matching, NeurIPS 2021 [[code]](https://github.com/XLearning-SCU/2021-NeurIPS-NCR)
MURAL: Multimodal, Multitask Retrieval Across Languages, arXiv 2021
Self-Supervised Learning from Web Data for Multimodal Retrieval, arXiv 2019
Look, Imagine and Match: Improving Textual-Visual Cross-Modal Retrieval with Generative Models, CVPR 2018
Scene-centric vs. Object-centric Image-Text Cross-modal Retrieval: A Reproducibility Study, ECIR 2023
Multimodal Co-learning
Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision, ICML 2021
Multimodal Co-learning: Challenges, Applications with Datasets, Recent Advances and Future Directions, arXiv 2021
Vokenization: Improving Language Understanding via Contextualized, Visually-Grounded Supervision, EMNLP 2020
Foundations of Multimodal Co-learning, Information Fusion 2020
Missing or Imperfect Modalities
A Variational Information Bottleneck Approach to Multi-Omics Data Integration, AISTATS 2021 [[code]](https://github.com/chl8856/DeepIMV)
SMIL: Multimodal Learning with Severely Missing Modality, AAAI 2021
Factorized Inference in Deep Markov Models for Incomplete Multimodal Time Series, arXiv 2019
Learning Representations from Imperfect Time Series Data via Tensor Rank Regularization, ACL 2019
Multimodal Deep Learning for Robust RGB-D Object Recognition, IROS 2015
Analysis of Multimodal Models
M2Lens: Visualizing and Explaining Multimodal Models for Sentiment Analysis, IEEE TVCG 2022
Decoupling the Role of Data, Attention, and Losses in Multimodal Transformers, TACL 2021
Does my multimodal model learn cross-modal interactions? It’s harder to tell than you might think!, EMNLP 2020
Blindfold Baselines for Embodied QA, NIPS 2018 Visually-Grounded Interaction and Language Workshop
Analyzing the Behavior of Visual Question Answering Models, EMNLP 2016
Knowledge Graphs and Knowledge Bases
MMKG: Multi-Modal Knowledge Graphs, ESWC 2019
Answering Visual-Relational Queries in Web-Extracted Knowledge Graphs, AKBC 2019
Embedding Multimodal Relational Data for Knowledge Base Completion, EMNLP 2018
A Multimodal Translation-Based Approach for Knowledge Graph Representation Learning, SEM 2018 [[code]](https://github.com/UKPLab/starsem18-multimodalKB)
Order-Embeddings of Images and Language, ICLR 2016 [[code]](https://github.com/ivendrov/order-embedding)
Building a Large-scale Multimodal Knowledge Base System for Answering Visual Queries, arXiv 2015
Intepretable Learning
Multimodal Explanations by Predicting Counterfactuality in Videos, CVPR 2019
Multimodal Explanations: Justifying Decisions and Pointing to the Evidence, CVPR 2018 [[code]](https://github.com/Seth-Park/MultimodalExplanations)
Do Explanations make VQA Models more Predictable to a Human?, EMNLP 2018
Towards Transparent AI Systems: Interpreting Visual Question Answering Models, ICML Workshop on Visualization for Deep Learning 2016
Generative Learning
MMVAE+: Enhancing the Generative Quality of Multimodal VAEs without Compromises, ICLR 2023 [[code]](https://github.com/epalu/mmvaeplus)
On the Limitations of Multimodal VAEs, ICLR 2022 [[code]](https://openreview.net/attachment?id=w-CPUXXrAj&name=supplementary_material)
Generalized Multimodal ELBO, ICLR 2021 [[code]](https://github.com/thomassutter/MoPoE)
Multimodal Generative Learning Utilizing Jensen-Shannon-Divergence, NeurIPS 2020 [[code]](https://github.com/thomassutter/mmjsd)
Self-supervised Disentanglement of Modality-specific and Shared Factors Improves Multimodal Generative Models, GCPR 2020 [[code]](https://github.com/imantdaunhawer/DMVAE)
Variational Mixture-of-Experts Autoencodersfor Multi-Modal Deep Generative Models, NeurIPS 2019 [[code]](https://github.com/iffsid/mmvae)
Few-shot Video-to-Video Synthesis, NeurIPS 2019 [[code]](https://nvlabs.github.io/few-shot-vid2vid/)
Multimodal Generative Models for Scalable Weakly-Supervised Learning, NeurIPS 2018 [[code1]](https://github.com/mhw32/multimodal-vae-public) [[code2]](https://github.com/panpan2/Multimo