Two Heads are Better than One: Simulating Large Transformers with Small Ones

NeurIPSSpotlight2025

Authors
Hantao Yu, Josh Alman
Venue
NeurIPS 2025
Track
Spotlight

TL;DR

The quadratic complexity of self‑attention prevents transformers from scaling effectively to long input sequences…

Opening excerpt from the authors’ abstract. source

Read the paper

Topics

transformer attention

← All NeurIPS 2025 Spotlight papers · Browse the whole archive