Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?

NeurIPSOral2025

Authors
Yang Yue, Zhiqi Chen, Rui Lu, Andrew Zhao, Zhaokai Wang, Yang Yue, Shiji Song, Gao Huang
Venue
NeurIPS 2025
Track
Oral

TL;DR

We systematically examine the current state of RLVR and surprisingly find that it does not elicit fundamentally new reasoning patterns—revealing a gap between the potential of RL and the actual impact of current RLVR methods.

Opening excerpt from the authors’ abstract. source

Read the paper

Topics

reinforcement learning reasoning llm

← All NeurIPS 2025 Oral papers · Browse the whole archive