Is it Thinking or Cheating? Detecting Implicit Reward Hacking by Measuring Reasoning Effort

ICLROral2026

Authors
Xinpeng Wang, Nitish Joshi, Barbara Plank, Rico Angell, He He
Affiliation
Ludwig-Maximilians-Universität München
Venue
ICLR 2026
Track
Oral

TL;DR

TRACE detects implicit reward hacking by measuring how quickly truncated reasoning suffices to pass verification, outperforming CoT monitoring and enabling hidden loopholes discovery.

Opening excerpt from the authors’ abstract. source

Read the paper

Topics

reasoning

← All ICLR 2026 Oral papers · Browse the whole archive