The Hawthorne Effect in Reasoning Models: Evaluating and Steering Test Awareness
NeurIPSSpotlight2025
TL;DR
Reasoning-focused LLMs sometimes alter their behavior when they detect that they are being evaluated—which can lead them to optimize for test-passing performance or to comply more readily with harmful…
Opening excerpt from the authors’ abstract. source
Read the paper
Topics
reasoning llm
← All NeurIPS 2025 Spotlight papers · Browse the whole archive