Reducing Belief Deviation in Reinforcement Learning for Active Reasoning of LLM Agents

ICLROral2026

Authors
Deyu Zou, Yongqiang Chen, Jianxiang Wang, Garry YANG, Mufei Li, Qing Da, James Cheng, Pan Li, Yu Gong
Affiliation
Department of Computer Science and Engineering, The Chinese University of Hong Kong
Venue
ICLR 2026
Track
Oral

TL;DR

Active reasoning requires large language model (LLM) agents to interact with external sources and strategically gather information to solve problems in multiple turns. Central to this process is belief tracking: maintaining an accurate representation of the underlying state and uncertainty in understanding and solving the problem.

Opening excerpt from the authors’ abstract. source

Read the paper

Topics

reinforcement learning large language model language model uncertainty reasoning agent llm

← All ICLR 2026 Oral papers · Browse the whole archive