Data Mixing Can Induce Phase Transitions in Knowledge Acquisition

NeurIPSSpotlight2025

Authors
Xinran Gu, Kaifeng Lyu, Jiazheng Li, Jingzhao Zhang
Venue
NeurIPS 2025
Track
Spotlight

TL;DR

Large Language Models (LLMs) are typically trained on data mixtures: most data come from web scrapes, while a small portion is curated from high-quality sources with dense domain-specific knowledge…

Opening excerpt from the authors’ abstract. source

Read the paper

Topics

large language model language model llm

← All NeurIPS 2025 Spotlight papers · Browse the whole archive