RAIN-Merging: A Gradient-Free Method to Enhance Instruction Following in Large Reasoning Models with Preserved Thinking Format

ICLROral2026

Authors
Zhehao Huang, Yuhang Liu, Baijiong Lin, Yixin Lou, Zhengbao He, Hanling Tian, Tao Li, Xiaolin Huang
Affiliation
Shanghai Jiaotong University
Venue
ICLR 2026
Track
Oral

TL;DR

Large reasoning models (LRMs) excel at a long chain of reasoning but often fail to faithfully follow instructions regarding output format, constraints, or specific requirements. We investigate whether this gap can be closed by integrating an instruction-tuned model (ITM) into an LRM.

Opening excerpt from the authors’ abstract. source

Read the paper

Topics

reasoning

← All ICLR 2026 Oral papers · Browse the whole archive