Vision Transformers Don't Need Trained Registers
NeurIPSSpotlight2025
TL;DR
We investigate the mechanism underlying a previously identified phenomenon in Vision Transformers -- the emergence of high-norm tokens that lead to noisy attention maps (Darcet et al…
Opening excerpt from the authors’ abstract. source
Read the paper
Topics
transformer attention
← All NeurIPS 2025 Spotlight papers · Browse the whole archive