Communication-Efficient Language Model Training Scales Reliably and Robustly: Scaling Laws for DiLoCo
NeurIPSSpotlight2025
TL;DR
As we scale to more massive machine learning models, the frequent synchronization demands inherent in data-parallel approaches create significant slowdowns, posing a critical challenge to further scal…
Opening excerpt from the authors’ abstract. source
Read the paper
Topics
language model scaling law efficient
← All NeurIPS 2025 Spotlight papers · Browse the whole archive