Language Models can Self-Improve at State-Value Estimation for Better Search

NeurIPSSpotlight2025

Authors
Ethan Mendes, Alan Ritter
Venue
NeurIPS 2025
Track
Spotlight

TL;DR

Collecting ground-truth rewards or human demonstrations for multi-step reasoning tasks is often prohibitively expensive, especially in interactive domains such as web tasks…

Opening excerpt from the authors’ abstract. source

Read the paper

Topics

language model reasoning

← All NeurIPS 2025 Spotlight papers · Browse the whole archive