MonarchAttention: Zero-Shot Conversion to Fast, Hardware-Aware Structured Attention
NeurIPSSpotlight2025
TL;DR
Transformers have achieved state-of-the-art performance across various tasks, but suffer from a notable quadratic complexity in sequence length due to the attention mechanism…
Opening excerpt from the authors’ abstract. source
Read the paper
Topics
transformer attention zero-shot
← All NeurIPS 2025 Spotlight papers · Browse the whole archive