Streaming Attention Approximation via Discrepancy Theory
NeurIPSSpotlight2025
TL;DR
Large language models (LLMs) have achieved impressive success, but their high memory requirements present challenges for long-context token generation…
Opening excerpt from the authors’ abstract. source
Read the paper
Topics
large language model language model attention memory theory llm
← All NeurIPS 2025 Spotlight papers · Browse the whole archive