Angular Steering: Behavior Control via Rotation in Activation Space
NeurIPSSpotlight2025
TL;DR
Controlling specific behaviors in large language models while preserving their general capabilities is a central challenge for safe and reliable artificial intelligence deployment…
Opening excerpt from the authors’ abstract. source
Read the paper
Topics
large language model language model control
← All NeurIPS 2025 Spotlight papers · Browse the whole archive