The Hawthorne Effect in Reasoning Models: Evaluating and Steering Test Awareness

NeurIPSSpotlight2025

Authors
Sahar Abdelnabi, Ahmed Salem
Venue
NeurIPS 2025
Track
Spotlight

TL;DR

Reasoning-focused LLMs sometimes alter their behavior when they detect that they are being evaluated—which can lead them to optimize for test-passing performance or to comply more readily with harmful…

Opening excerpt from the authors’ abstract. source

Read the paper

Topics

reasoning llm

← All NeurIPS 2025 Spotlight papers · Browse the whole archive