Twilight: Adaptive Attention Sparsity with Hierarchical Top-$p$ Pruning
NeurIPSSpotlight2025
TL;DR
Leveraging attention sparsity to accelerate long-context large language models (LLMs) has been of great importance recently…
Opening excerpt from the authors’ abstract. source
Read the paper
Topics
large language model language model attention sparsity llm rag
← All NeurIPS 2025 Spotlight papers · Browse the whole archive