Twilight: Adaptive Attention Sparsity with Hierarchical Top-$p$ Pruning

NeurIPSSpotlight2025

Authors
Chaofan Lin, Jiaming Tang, Shuo Yang, Hanshuo Wang, Tian Tang, Boyu Tian, Ion Stoica, Song Han, Mingyu Gao
Venue
NeurIPS 2025
Track
Spotlight

TL;DR

Leveraging attention sparsity to accelerate long-context large language models (LLMs) has been of great importance recently…

Opening excerpt from the authors’ abstract. source

Read the paper

Topics

large language model language model attention sparsity llm rag

← All NeurIPS 2025 Spotlight papers · Browse the whole archive