Sampling-Efficient Test-Time Scaling: Self-Estimating the Best-of-N Sampling in Early Decoding
NeurIPSSpotlight2025
TL;DR
Test-time scaling enhances large language model performance by allocating additional compute resources during decoding…
Opening excerpt from the authors’ abstract. source
Read the paper
Topics
large language model language model efficient
← All NeurIPS 2025 Spotlight papers · Browse the whole archive