Head Pursuit: Probing Attention Specialization in Multimodal Transformers

NeurIPSSpotlight2025

Authors
Lorenzo Basile, Valentino Maiorca, Diego Doimo, Francesco Locatello, Alberto Cazzaniga
Venue
NeurIPS 2025
Track
Spotlight

TL;DR

Language and vision-language models have shown impressive performance across a wide range of tasks, but their internal mechanisms remain only partly understood…

Opening excerpt from the authors’ abstract. source

Read the paper

Topics

vision-language language model transformer multimodal attention

← All NeurIPS 2025 Spotlight papers · Browse the whole archive