Critical Batch Size Revisited: A Simple Empirical Approach to Large-Batch Language Model Training

NeurIPSSpotlight2025

Authors
William Merrill, Shane Arora, Dirk Groeneveld, Hannaneh Hajishirzi
Venue
NeurIPS 2025
Track
Spotlight

TL;DR

The right batch size is important when training language models at scale: a large batch size is necessary for fast training, but a batch size that is *too large* will harm token efficiency…

Opening excerpt from the authors’ abstract. source

Read the paper

Topics

language model

← All NeurIPS 2025 Spotlight papers · Browse the whole archive