Why Do Some Language Models Fake Alignment While Others Don't?
NeurIPSSpotlight2025
TL;DR
*Alignment Faking in Large Language Models* presented a demonstration of Claude 3 Opus and Claude 3…
Opening excerpt from the authors’ abstract. source
Read the paper
Topics
large language model language model alignment
← All NeurIPS 2025 Spotlight papers · Browse the whole archive