Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free

NeurIPSOral2025

Authors
Zihan Qiu, Zekun Wang, Bo Zheng, Zeyu Huang, Kaiyue Wen, Songlin Yang, Rui Men, Le Yu, Fei Huang, Suozhi Huang, Dayiheng Liu, Jingren Zhou, Junyang Lin
Affiliation
Alibaba Group
Venue
NeurIPS 2025
Track
Oral

TL;DR

We find applying a query-dependent head-specific sigmoid gate after the Scaled Dot-Product Attention (SDPA) consistently improves performance, improves scaling properties and mitigates the `massive activation' and `attention sink'.

Opening excerpt from the authors’ abstract. source

Read the paper

Topics

large language model language model attention sparsity

← All NeurIPS 2025 Oral papers · Browse the whole archive