Language Models can Self-Improve at State-Value Estimation for Better Search
NeurIPSSpotlight2025
TL;DR
Collecting ground-truth rewards or human demonstrations for multi-step reasoning tasks is often prohibitively expensive, especially in interactive domains such as web tasks…
Opening excerpt from the authors’ abstract. source
Read the paper
Topics
language model reasoning
← All NeurIPS 2025 Spotlight papers · Browse the whole archive