From Condensation to Rank Collapse: A Two-Stage Analysis of Transformer Training Dynamics

NeurIPSOral2025

Authors
Zheng-An Chen, Tao Luo
Affiliation
Shanghai Jiaotong University
Venue
NeurIPS 2025
Track
Oral

TL;DR

Although transformer-based models have shown exceptional empirical performance, the fundamental principles governing their training dynamics are inadequately characterized beyond configuration-specific studies. Inspired by empirical evidence showing improved reasoning capabilities under small initialization scales in language models, we employ the…

Opening excerpt from the authors’ abstract. source

Read the paper

Topics

language model transformer reasoning

← All NeurIPS 2025 Oral papers · Browse the whole archive