Absolute Zero: Reinforced Self-play Reasoning with Zero Data
NeurIPSSpotlight2025
TL;DR
Reinforcement learning with verifiable rewards (RLVR) has shown promise in enhancing the reasoning capabilities of large language models by learning directly from rule-based outcome rewards…
Opening excerpt from the authors’ abstract. source
Read the paper
Topics
reinforcement learning large language model language model reasoning
← All NeurIPS 2025 Spotlight papers · Browse the whole archive