OPUS: Towards Efficient and Principled Data Selection in Large Language Model Pre-training in Every Iteration

ICMLOral2026

Authors
Shaobo Wang, Xuan Ouyang, Tianyi Xu, Yuzheng Hu, Jialin Liu, Guo Chen, Tianyu Zhang, Junhao Zheng, Kexin Yang, Xingzhang Ren, Dayiheng Liu, Linfeng Zhang
Venue
ICML 2026
Track
Oral

TL;DR

No summary has been collected for this paper yet. Read the abstract at the authoritative source below.

Read the paper

Topics

large language model language model pre-training efficient

← All ICML 2026 Oral papers · Browse the whole archive