TROLL: Trust Regions Improve Reinforcement Learning for Large Language Models
ICLROral2026
TL;DR
Replacing PPO's clipping objective with more principled trust regions improves RL from verifiable rewards.
Opening excerpt from the authors’ abstract. source
Read the paper
Topics
reinforcement learning large language model language model