On the Value of Cross-Modal Misalignment in Multimodal Representation Learning
NeurIPSSpotlight2025
TL;DR
Multimodal representation learning, exemplified by multimodal contrastive learning (MMCL) using image-text pairs, aims to learn powerful representations by aligning cues across modalities…
Opening excerpt from the authors’ abstract. source
Read the paper
Topics
representation learning contrastive multimodal alignment
← All NeurIPS 2025 Spotlight papers · Browse the whole archive