WAVE: Learning Unified & Versatile Audio-Visual Embeddings with Multimodal LLM
ICLROral2026
TL;DR
This paper builds a versatile audio-visual embedding LLM, which can not only achieve any-to-any retrieval but also generate prompt-aware embeddings.
Opening excerpt from the authors’ abstract. source
Read the paper
Topics
multimodal retrieval audio llm