From Condensation to Rank Collapse: A Two-Stage Analysis of Transformer Training Dynamics
NeurIPSOral2025
TL;DR
Although transformer-based models have shown exceptional empirical performance, the fundamental principles governing their training dynamics are inadequately characterized beyond configuration-specific studies. Inspired by empirical evidence showing improved reasoning capabilities under small initialization scales in language models, we employ the…
Opening excerpt from the authors’ abstract. source
Read the paper
Topics
language model transformer reasoning