ZeroS: Zero‑Sum Linear Attention for Efficient Transformers
NeurIPSSpotlight2025
TL;DR
Linear attention methods offer Transformers $O(N)$ complexity but typically underperform standard softmax attention…
Opening excerpt from the authors’ abstract. source
Read the paper
Topics
transformer attention efficient
← All NeurIPS 2025 Spotlight papers · Browse the whole archive