TROLL: Trust Regions Improve Reinforcement Learning for Large Language Models

ICLROral2026

Authors
Philipp Becker, Niklas Freymuth, Serge Thilges, Fabian Otto, Gerhard Neumann
Affiliation
Facebook
Venue
ICLR 2026
Track
Oral

TL;DR

Replacing PPO's clipping objective with more principled trust regions improves RL from verifiable rewards.

Opening excerpt from the authors’ abstract. source

Read the paper

Topics

reinforcement learning large language model language model

← All ICLR 2026 Oral papers · Browse the whole archive