Head Pursuit: Probing Attention Specialization in Multimodal Transformers
NeurIPSSpotlight2025
TL;DR
Language and vision-language models have shown impressive performance across a wide range of tasks, but their internal mechanisms remain only partly understood…
Opening excerpt from the authors’ abstract. source
Read the paper
Topics
vision-language language model transformer multimodal attention
← All NeurIPS 2025 Spotlight papers · Browse the whole archive