Two Heads are Better than One: Simulating Large Transformers with Small Ones
NeurIPSSpotlight2025
TL;DR
The quadratic complexity of self‑attention prevents transformers from scaling effectively to long input sequences…
Opening excerpt from the authors’ abstract. source
Read the paper
Topics
transformer attention
← All NeurIPS 2025 Spotlight papers · Browse the whole archive