Multimodal Aligned Semantic Knowledge for Unpaired Image-text Matching
ICLROral2026
TL;DR
We propose multimodal aligned semantic knowledge, which leverages word embeddings as bridges to associate words with prototypes, capturing semantic relationships between words and further utilizing information from OOD words.
Opening excerpt from the authors’ abstract. source
Read the paper
Topics
multimodal rag