Inference-Time Reward Hacking in Large Language Models

NeurIPSSpotlight2025

Authors
Hadi Khalaf, Claudio Mayrink Verdun, Alex Oesterling, Himabindu Lakkaraju, Flavio Calmon
Venue
NeurIPS 2025
Track
Spotlight

TL;DR

A common paradigm to improve the performance of large language models is optimizing for a reward model…

Opening excerpt from the authors’ abstract. source

Read the paper

Topics

large language model language model

← All NeurIPS 2025 Spotlight papers · Browse the whole archive