Sampling-Efficient Test-Time Scaling: Self-Estimating the Best-of-N Sampling in Early Decoding

NeurIPSSpotlight2025

Authors
Yiming Wang, Pei Zhang, Siyuan Huang, Baosong Yang, Zhuosheng Zhang, Fei Huang, Rui Wang
Venue
NeurIPS 2025
Track
Spotlight

TL;DR

Test-time scaling enhances large language model performance by allocating additional compute resources during decoding…

Opening excerpt from the authors’ abstract. source

Read the paper

Topics

large language model language model efficient

← All NeurIPS 2025 Spotlight papers · Browse the whole archive