MetaEmbed: Scaling Multimodal Retrieval at Test-Time with Flexible Late Interaction

ICLROral2026

Authors
Zilin Xiao, Qi Ma, Mengting Gu, Chun-cheng Jason Chen, Xintao Chen, Vicente Ordonez, Vijai Mohan
Affiliation
Rice University
Venue
ICLR 2026
Track
Oral

TL;DR

Universal multimodal embedding models have achieved great success in capturing semantic relevance between queries and candidates. However, current methods either condense queries and candidates into a single vector, potentially limiting the expressiveness for fine-grained information, or produce too many vectors that are prohibitively expensive...

Opening excerpt from the authors’ abstract. source

Read the paper

Topics

multimodal retrieval

← All ICLR 2026 Oral papers · Browse the whole archive