Why Low-Precision Transformer Training Fails: An Analysis on Flash Attention

ICLROral2026

Authors
Haiquan Qiu, Quanming Yao
Affiliation
Tsinghua University, Tsinghua University
Venue
ICLR 2026
Track
Oral

TL;DR

For the first time, we mechanistically explain why low-precision training with flash attention fails, identifying a vicious cycle of rounding errors and proposing a simple, effective fix.

Opening excerpt from the authors’ abstract. source

Read the paper

Topics

transformer attention

← All ICLR 2026 Oral papers · Browse the whole archive