pliang279/awesome-multimodal-ml

★ 6,931⑂ 0

Reading list for research topics in multimodal machine learning

About pliang279/awesome-multimodal-ml

pliang279/awesome-multimodal-ml is an open-source project on GitHub, mainly written in several languages. Reading list for research topics in multimodal machine learning It currently holds 6,931 stars and 0 forks with 0 open issues, and was last pushed on an unknown date (repository created unknown).

Project Overview

AI Homed tracks it on the AI Image Projects board and on the AI AI Image Projects list.

GitHub Repository Details

Repository pliang279/awesome-multimodal-ml · default branch - · size 0 KB · watchers 0 · source: GitHub REST API and repository README

README

Awesome Multimodal Machine Learning

By Paul Liang (pliang@cs.cmu.edu), Machine Learning Department and Language Technologies Institute, CMU, with help from members of the MultiComp Lab at LTI, CMU. If there are any areas, papers, and datasets I missed, please let me know!

Course content + workshops

Check out our comprehsensive tutorial paper Foundations and Recent Trends in Multimodal Machine Learning: Principles, Challenges, and Open Questions.

Tutorials on Multimodal Machine Learning at CVPR 2022 and NAACL 2022, slides and videos here.

New course 11-877 Advanced Topics in Multimodal Machine Learning Spring 2022 @ CMU. It will primarily be reading and discussion-based. We plan to post discussion probes, relevant papers, and summarized discussion highlights every week on the website.

Public course content and lecture videos from 11-777 Multimodal Machine Learning, Fall 2020 @ CMU.

Table of Contents

Research Papers

Survey Papers

Foundations and Trends in Multimodal Machine Learning: Principles, Challenges, and Open Questions, arxiv 2023

Multimodal Learning with Transformers: A Survey, TPAMI 2023

Trends in Integration of Vision and Language Research: A Survey of Tasks, Datasets, and Methods, JAIR 2021

Experience Grounds Language, EMNLP 2020

A Survey of Reinforcement Learning Informed by Natural Language, IJCAI 2019

Multimodal Machine Learning: A Survey and Taxonomy, TPAMI 2019

Multimodal Intelligence: Representation Learning, Information Fusion, and Applications, arXiv 2019

Deep Multimodal Representation Learning: A Survey, arXiv 2019

Guest Editorial: Image and Language Understanding, IJCV 2017

Representation Learning: A Review and New Perspectives, TPAMI 2013

A Survey of Socially Interactive Robots, 2003

Core Areas

Multimodal Representations

Identifiability Results for Multimodal Contrastive Learning, ICLR 2023 [[code]](https://github.com/imantdaunhawer/multimodal-contrastive-learning)

Unpaired Vision-Language Pre-training via Cross-Modal CutMix, ICML 2022.

Balanced Multimodal Learning via On-the-fly Gradient Modulation, CVPR 2022

Unsupervised Voice-Face Representation Learning by Cross-Modal Prototype Contrast, IJCAI 2021 [[code]](https://github.com/Cocoxili/CMPC)

Towards a Unified Foundation Model: Jointly Pre-Training Transformers on Unpaired Images and Text, arXiv 2021

FLAVA: A Foundational Language And Vision Alignment Model, arXiv 2021

Transformer is All You Need: Multimodal Multitask Learning with a Unified Transformer, arXiv 2021

MultiBench: Multiscale Benchmarks for Multimodal Representation Learning, NeurIPS 2021 [[code]](https://github.com/pliang279/MultiBench)

Perceiver: General Perception with Iterative Attention, ICML 2021 [[code]](https://github.com/deepmind/deepmind-research/tree/master/perceiver)

Learning Transferable Visual Models From Natural Language Supervision, arXiv 2021 [[blog]](blog) [[code]](https://github.com/OpenAI/CLIP)

VinVL: Revisiting Visual Representations in Vision-Language Models, arXiv 2021 [[blog]](https://www.microsoft.com/en-us/research/blog/vinvl-advancing-the-state-of-the-art-for-vision-language-models/?OCID=msr_blog_VinVL_fb) [[code]](https://github.com/pzzhang/VinVL)

Learning Transferable Visual Models From Natural Language Supervision, arXiv 2020 [[blog]](https://openai.com/blog/clip/) [[code]](https://github.com/openai/CLIP)

12-in-1: Multi-Task Vision and Language Representation Learning, CVPR 2020 [[code]](https://github.com/facebookresearch/vilbert-multi-task)

Watching the World Go By: Representation Learning from Unlabeled Videos, arXiv 2020

Learning Video Representations using Contrastive Bidirectional Transformer, arXiv 2019

Visual Concept-Metaconcept Learning, NeurIPS 2019 [[code]](http://vcml.csail.mit.edu/)

OmniNet: A Unified Architecture for Multi-modal Multi-task Learning, arXiv 2019 [[code]](https://github.com/subho406/OmniNet)

Learning Representations by Maximizing Mutual Information Across Views, arXiv 2019 [[code]](https://github.com/Philip-Bachman/amdim-public)

ViCo: Word Embeddings from Visual Co-occurrences, ICCV 2019 [[code]](https://github.com/BigRedT/vico)

Unified Visual-Semantic Embeddings: Bridging Vision and Language With Structured Meaning Representations, CVPR 2019

Multi-Task Learning of Hierarchical Vision-Language Representation, CVPR 2019

Learning Factorized Multimodal Representations, ICLR 2019 [[code]](https://github.com/pliang279/factorized/)

A Probabilistic Framework for Multi-view Feature Learning with Many-to-many Associations via Neural Networks, ICML 2018

Do Neural Network Cross-Modal Mappings Really Bridge Modalities?, ACL 2018

Learning Robust Visual-Semantic Embeddings, ICCV 2017

Deep Multimodal Representation Learning from Temporal Data, CVPR 2017

Is an Image Worth More than a Thousand Words? On the Fine-Grain Semantic Differences between Visual and Linguistic Representations, COLING 2016

Combining Language and Vision with a Multimodal Skip-gram Model, NAACL 2015

Deep Fragment Embeddings for Bidirectional Image Sentence Mapping, NIPS 2014

Multimodal Learning with Deep Boltzmann Machines, JMLR 2014

Learning Grounded Meaning Representations with Autoencoders, ACL 2014

DeViSE: A Deep Visual-Semantic Embedding Model, NeurIPS 2013

Multimodal Deep Learning, ICML 2011

Multimodal Fusion

Robust Contrastive Learning against Noisy Views, arXiv 2022

Cooperative Learning for Multi-view Analysis, arXiv 2022

What Makes Multi-modal Learning Better than Single (Provably), NeurIPS 2021

Efficient Multi-Modal Fusion with Diversity Analysis, ACMMM 2021

Attention Bottlenecks for Multimodal Fusion, NeurIPS 2021

VMLoc: Variational Fusion For Learning-Based Multimodal Camera Localization, AAAI 2021

Trusted Multi-View Classification, ICLR 2021 [[code]](https://github.com/hanmenghan/TMC)

Deep-HOSeq: Deep Higher-Order Sequence Fusion for Multimodal Sentiment Analysis, ICDM 2020

Removing Bias in Multi-modal Classifiers: Regularization by Maximizing Functional Entropies, NeurIPS 2020 [[code]](https://github.com/itaigat/removing-bias-in-multi-modal-classifiers)

Deep Multimodal Fusion by Channel Exchanging, NeurIPS 2020 [[code]](https://github.com/yikaiw/CEN)

What Makes Training Multi-Modal Classification Networks Hard?, CVPR 2020

Dynamic Fusion for Multimodal Data, arXiv 2019

DeepCU: Integrating Both Common and Unique Latent Information for Multimodal Sentiment Analysis, IJCAI 2019 [[code]](https://github.com/sverma88/DeepCU-IJCAI19)

Deep Multimodal Multilinear Fusion with High-order Polynomial Pooling, NeurIPS 2019

XFlow: Cross-modal Deep Neural Networks for Audiovisual Classification, IEEE TNNLS 2019 [[code]](https://github.com/catalina17/XFlow)

MFAS: Multimodal Fusion Architecture Search, CVPR 2019

The Neuro-Symbolic Concept Learner: Interpreting Scenes, Words, and Sentences From Natural Supervision, ICLR 2019 [[code]](http://nscl.csail.mit.edu/)

Unifying and merging well-trained deep neural networks for inference stage, IJCAI 2018 [[code]](https://github.com/ivclab/NeuralMerger)

Efficient Low-rank Multimodal Fusion with Modality-Specific Factors, ACL 2018 [[code]](https://github.com/Justin1904/Low-rank-Multimodal-Fusion)

Memory Fusion Network for Multi-view Sequential Learning, AAAI 2018 [[code]](https://github.com/pliang279/MFN)

Tensor Fusion Network for Multimodal Sentiment Analysis, EMNLP 2017 [[code]](https://github.com/A2Zadeh/TensorFusionNetwork)

Jointly Modeling Deep Video and Compositional Text to Bridge Vision and Language in a Unified Framework, AAAI 2015

A co-regularized approach to semi-supervised learning with multiple views, ICML 2005

Multimodal Alignment

Reconsidering Representation Alignment for Multi-view Clustering, CVPR 2021 [[code]](https://github.com/DanielTrosten/mvc)

CoMIR: Contrastive Multimodal Image Representation for Registration, NeurIPS 2020 [[code]](https://github.com/MIDA-group/CoMIR)

Multimodal Transformer for Unaligned Multimodal Language Sequences, ACL 2019 [[code]](https://github.com/yaohungt/Multimodal-Transformer)

Temporal Cycle-Consistency Learning, CVPR 2019 [[code]](https://github.com/google-research/google-research/tree/master/tcc)

See, Hear, and Read: Deep Aligned Representations, arXiv 2017

On Deep Multi-View Representation Learning, ICML 2015

Unsupervised Alignment of Natural Language Instructions with Video Segments, AAAI 2014

Multimodal Alignment of Videos, MM 2014

Deep Canonical Correlation Analysis, ICML 2013 [[code]](https://github.com/VahidooX/DeepCCA)

Multimodal Pretraining

Align before Fuse: Vision and Language Representation Learning with Momentum Distillation, NeurIPS 2021 Spotlight [[code]](https://github.com/salesforce/ALBEF)

Less is More: ClipBERT for Video-and-Language Learning via Sparse Sampling, CVPR 2021 [[code]](https://github.com/jayleicn/ClipBERT)

Transformer is All You Need: Multimodal Multitask Learning with a Unified Transformer, arXiv 2021

Large-Scale Adversarial Training for Vision-and-Language Representation Learning, NeurIPS 2020 [[code]](https://github.com/zhegan27/VILLA)

Vokenization: Improving Language Understanding with Contextualized, Visual-Grounded Supervision, EMNLP 2020 [[code]](https://github.com/airsplay/vokenization)

Integrating Multimodal Information in Large Pretrained Transformers, ACL 2020

VL-BERT: Pre-training of Generic Visual-Linguistic Representations, arXiv 2019 [[code]](https://github.com/jackroos/VL-BERT)

VisualBERT: A Simple and Performant Baseline for Vision and Language, arXiv 2019 [[code]](https://github.com/uclanlp/visualbert)

ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks, NeurIPS 2019 [[code]](https://github.com/jiasenlu/vilbert_beta)

Unicoder-VL: A Universal Encoder for Vision and Language by Cross-modal Pre-training, arXiv 2019

LXMERT: Learning Cross-Modality Encoder Representations from Transformers, EMNLP 2019 [[code]](https://github.com/airsplay/lxmert)

VideoBERT: A Joint Model for Video and Language Representation Learning, ICCV 2019

Multimodal Translation

Zero-Shot Text-to-Image Generation, ICML 2021 [[code]](https://github.com/openai/DALL-E)

Translate-to-Recognize Networks for RGB-D Scene Recognition, CVPR 2019 [[code]](https://github.com/ownstyledu/Translate-to-Recognize-Networks)

Language2Pose: Natural Language Grounded Pose Forecasting, 3DV 2019 [[code]](http://chahuja.com/language2pose/)

Reconstructing Faces from Voices, NeurIPS 2019 [[code]](https://github.com/cmu-mlsp/reconstructing_faces_from_voices)

Speech2Face: Learning the Face Behind a Voice, CVPR 2019 [[code]](https://speech2face.github.io/)

Found in Translation: Learning Robust Joint Representations by Cyclic Translations Between Modalities, AAAI 2019 [[code]](https://github.com/hainow/MCTN)

Natural TTS Synthesis by Conditioning Wavenet on Mel Spectrogram Predictions, ICASSP 2018 [[code]](https://github.com/NVIDIA/tacotron2)

Crossmodal Retrieval

Learning with Noisy Correspondence for Cross-modal Matching, NeurIPS 2021 [[code]](https://github.com/XLearning-SCU/2021-NeurIPS-NCR)

MURAL: Multimodal, Multitask Retrieval Across Languages, arXiv 2021

Self-Supervised Learning from Web Data for Multimodal Retrieval, arXiv 2019

Look, Imagine and Match: Improving Textual-Visual Cross-Modal Retrieval with Generative Models, CVPR 2018

Scene-centric vs. Object-centric Image-Text Cross-modal Retrieval: A Reproducibility Study, ECIR 2023

Multimodal Co-learning

Self-Supervised Learning in Event Sequences: A Comparative Study and Hybrid Approach of Generative Modeling and Contrastive Learning, arXiv 2024

Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision, ICML 2021

Multimodal Co-learning: Challenges, Applications with Datasets, Recent Advances and Future Directions, arXiv 2021

Vokenization: Improving Language Understanding via Contextualized, Visually-Grounded Supervision, EMNLP 2020

Foundations of Multimodal Co-learning, Information Fusion 2020

Missing or Imperfect Modalities

A Variational Information Bottleneck Approach to Multi-Omics Data Integration, AISTATS 2021 [[code]](https://github.com/chl8856/DeepIMV)

SMIL: Multimodal Learning with Severely Missing Modality, AAAI 2021

Factorized Inference in Deep Markov Models for Incomplete Multimodal Time Series, arXiv 2019

Learning Representations from Imperfect Time Series Data via Tensor Rank Regularization, ACL 2019

Multimodal Deep Learning for Robust RGB-D Object Recognition, IROS 2015

Analysis of Multimodal Models

M2Lens: Visualizing and Explaining Multimodal Models for Sentiment Analysis, IEEE TVCG 2022

Decoupling the Role of Data, Attention, and Losses in Multimodal Transformers, TACL 2021

Does my multimodal model learn cross-modal interactions? It’s harder to tell than you might think!, EMNLP 2020

Blindfold Baselines for Embodied QA, NIPS 2018 Visually-Grounded Interaction and Language Workshop

Analyzing the Behavior of Visual Question Answering Models, EMNLP 2016

Knowledge Graphs and Knowledge Bases

MMKG: Multi-Modal Knowledge Graphs, ESWC 2019

Answering Visual-Relational Queries in Web-Extracted Knowledge Graphs, AKBC 2019

Embedding Multimodal Relational Data for Knowledge Base Completion, EMNLP 2018

A Multimodal Translation-Based Approach for Knowledge Graph Representation Learning, SEM 2018 [[code]](https://github.com/UKPLab/starsem18-multimodalKB)

Order-Embeddings of Images and Language, ICLR 2016 [[code]](https://github.com/ivendrov/order-embedding)

Building a Large-scale Multimodal Knowledge Base System for Answering Visual Queries, arXiv 2015

Intepretable Learning

Multimodal Explanations by Predicting Counterfactuality in Videos, CVPR 2019

Multimodal Explanations: Justifying Decisions and Pointing to the Evidence, CVPR 2018 [[code]](https://github.com/Seth-Park/MultimodalExplanations)

Do Explanations make VQA Models more Predictable to a Human?, EMNLP 2018

Towards Transparent AI Systems: Interpreting Visual Question Answering Models, ICML Workshop on Visualization for Deep Learning 2016

Generative Learning

MMVAE+: Enhancing the Generative Quality of Multimodal VAEs without Compromises, ICLR 2023 [[code]](https://github.com/epalu/mmvaeplus)

On the Limitations of Multimodal VAEs, ICLR 2022 [[code]](https://openreview.net/attachment?id=w-CPUXXrAj&name=supplementary_material)

Generalized Multimodal ELBO, ICLR 2021 [[code]](https://github.com/thomassutter/MoPoE)

Multimodal Generative Learning Utilizing Jensen-Shannon-Divergence, NeurIPS 2020 [[code]](https://github.com/thomassutter/mmjsd)

Self-supervised Disentanglement of Modality-specific and Shared Factors Improves Multimodal Generative Models, GCPR 2020 [[code]](https://github.com/imantdaunhawer/DMVAE)

Variational Mixture-of-Experts Autoencodersfor Multi-Modal Deep Generative Models, NeurIPS 2019 [[code]](https://github.com/iffsid/mmvae)

Few-shot Video-to-Video Synthesis, NeurIPS 2019 [[code]](https://nvlabs.github.io/few-shot-vid2vid/)

Multimodal Generative Models for Scalable Weakly-Supervised Learning, NeurIPS 2018 [[code1]](https://github.com/mhw32/multimodal-vae-public) [[code2]](https://github.com/panpan2/Multimo

GitHub Stars & Activity

6,931Stars
0Forks
0Open issues
-Language

GitHub Popularity

GitHub stars6,931
Forks0
Open issues0
Primary language-
License-
Stars gained today0
Created-
Last pushed-

Trending History

Trending statusnot on today's boards

Related AI Projects

1

opencv / opencv

C++★ 90,853⑂ 0
2
3

d2l-ai / d2l-zh

Python★ 80,716⑂ 0
4

microsoft / AI-For-Beginners

Jupyter Notebook★ 68,541⑂ 0
5

ultralytics / ultralytics

Python★ 61,651⑂ 0
6

ultralytics / yolov5

Python★ 58,015⑂ 0
7
8

roboflow / supervision

Python★ 50,364⑂ 0

More AI Rankings