Pre-training under infinite compute
ICLROral2026
TL;DR
Since compute grows faster than the web, we design simple recipes that improve the asymptote of compute scaling laws to be 5x data efficient, offering better performance with sufficient compute.
Opening excerpt from the authors’ abstract. source
Read the paper
Topics
pre-training scaling law efficient