Taming Momentum: Rethinking Optimizer States Through Low-Rank Approximation

ICLROral2026

Authors
Zhengbo Wang, Jian Liang, Ran He, Zilei Wang, Tieniu Tan
Affiliation
University of Science and Technology of China
Venue
ICLR 2026
Track
Oral

TL;DR

Modern optimizers like Adam and Muon are central to training large language models, but their reliance on first- and second-order momenta introduces significant memory overhead, which constrains scalability and computational efficiency. In this work, we reframe the exponential moving average (EMA) used in these momenta as the training of a linea...

Opening excerpt from the authors’ abstract. source

Read the paper

Topics

large language model language model memory rag

← All ICLR 2026 Oral papers · Browse the whole archive