Streaming Attention Approximation via Discrepancy Theory

NeurIPSSpotlight2025

Authors
Ekaterina Kochetkova, Kshiteej Sheth, Insu Han, Amir Zandieh, Michael Kapralov
Venue
NeurIPS 2025
Track
Spotlight

TL;DR

Large language models (LLMs) have achieved impressive success, but their high memory requirements present challenges for long-context token generation…

Opening excerpt from the authors’ abstract. source

Read the paper

Topics

large language model language model attention memory theory llm

← All NeurIPS 2025 Spotlight papers · Browse the whole archive