Taming Momentum: Rethinking Optimizer States Through Low-Rank Approximation
ICLROral2026
TL;DR
Modern optimizers like Adam and Muon are central to training large language models, but their reliance on first- and second-order momenta introduces significant memory overhead, which constrains scalability and computational efficiency. In this work, we reframe the exponential moving average (EMA) used in these momenta as the training of a linea...
Opening excerpt from the authors’ abstract. source
Read the paper
Topics
large language model language model memory rag