Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free
NeurIPSOral2025
TL;DR
We find applying a query-dependent head-specific sigmoid gate after the Scaled Dot-Product Attention (SDPA) consistently improves performance, improves scaling properties and mitigates the `massive activation' and `attention sink'.
Opening excerpt from the authors’ abstract. source
Read the paper
Topics
large language model language model attention sparsity