How Learning Rate Decay Wastes Your Best Data in Curriculum-Based LLM Pretraining

ICLROral2026

Authors
Kairong Luo, Zhenbo Sun, Haodong Wen, Xinyu Shi, Jiarui Cui, Chenyi Dang, Kaifeng Lyu, Wenguang Chen
Affiliation
Tsinghua University
Venue
ICLR 2026
Track
Oral

TL;DR

Use model weight average to enhance curriculum learning in LLM pretraining.

Opening excerpt from the authors’ abstract. source

Read the paper

Topics

pretraining llm rag

← All ICLR 2026 Oral papers · Browse the whole archive