Right Question is Already Half the Answer: Fully Unsupervised LLM Reasoning Incentivization

NeurIPSSpotlight2025

Authors
Qingyang Zhang, Haitao Wu, Changqing Zhang, Peilin Zhao, Yatao Bian
Venue
NeurIPS 2025
Track
Spotlight

TL;DR

Existing methods to enhance the reasoning capability of large language models predominantly rely on supervised fine-tuning (SFT) followed by reinforcement learning (RL) on reasoning-specific data…

Opening excerpt from the authors’ abstract. source

Read the paper

Topics

reinforcement learning large language model language model fine-tuning reasoning llm

← All NeurIPS 2025 Spotlight papers · Browse the whole archive