The Art of Scaling Reinforcement Learning Compute for LLMs

ICLROral2026

Authors
Fnu Devvrit, Lovish Madaan, Rishabh Tiwari, Rachit Bansal, Sai Surya Duvvuri, Manzil Zaheer, Inderjit S Dhillon, David Brandfonbrener, Rishabh Agarwal
Affiliation
, University of Texas at Austin
Venue
ICLR 2026
Track
Oral

TL;DR

We study compute scaling properties of RL methods on LLMs…

Opening excerpt from the authors’ abstract. source

Read the paper

Topics

reinforcement learning llm

← All ICLR 2026 Oral papers · Browse the whole archive