Reducing Belief Deviation in Reinforcement Learning for Active Reasoning of LLM Agents
ICLROral2026
TL;DR
Active reasoning requires large language model (LLM) agents to interact with external sources and strategically gather information to solve problems in multiple turns. Central to this process is belief tracking: maintaining an accurate representation of the underlying state and uncertainty in understanding and solving the problem.
Opening excerpt from the authors’ abstract. source
Read the paper
Topics
reinforcement learning large language model language model uncertainty reasoning agent llm