Activation Control for Efficiently Eliciting Long Chain-of-thought Ability of Language Models
NeurIPSSpotlight2025
TL;DR
Despite the remarkable reasoning performance, eliciting the long chain-of-thought(CoT) ability in large language models(LLMs) typically requires costly reinforcement learning or supervised fine-tuning…
Opening excerpt from the authors’ abstract. source
Read the paper
Topics
reinforcement learning large language model chain-of-thought language model fine-tuning efficient reasoning control llm
← All NeurIPS 2025 Spotlight papers · Browse the whole archive