WAVE: Learning Unified & Versatile Audio-Visual Embeddings with Multimodal LLM

ICLROral2026

Authors
Changli Tang, Qinfan Xiao, Ke Mei, Tianyi Wang, Fengyun Rao, Chao Zhang
Affiliation
Tsinghua University, Tsinghua University
Venue
ICLR 2026
Track
Oral

TL;DR

This paper builds a versatile audio-visual embedding LLM, which can not only achieve any-to-any retrieval but also generate prompt-aware embeddings.

Opening excerpt from the authors’ abstract. source

Read the paper

Topics

multimodal retrieval audio llm

← All ICLR 2026 Oral papers · Browse the whole archive