Improved Representation Steering for Language Models
NeurIPSSpotlight2025
TL;DR
Steering methods for language models (LMs) seek to provide fine-grained and interpretable control over model generations by variously changing model inputs, weights, or representations to adjust behav…
Opening excerpt from the authors’ abstract. source
Read the paper
Topics
language model control
← All NeurIPS 2025 Spotlight papers · Browse the whole archive