Inference-Time Reward Hacking in Large Language Models
NeurIPSSpotlight2025
TL;DR
A common paradigm to improve the performance of large language models is optimizing for a reward model…
Opening excerpt from the authors’ abstract. source
Read the paper
Topics
large language model language model
← All NeurIPS 2025 Spotlight papers · Browse the whole archive