LoongRL: Reinforcement Learning for Advanced Reasoning over Long Contexts
ICLROral2026
TL;DR
Reasoning over long contexts is essential for large language models. While reinforcement learning (RL) enhances short-context reasoning by inducing "Aha" moments in chain-of-thought, the advanced thinking patterns required for long-context reasoning remain largely unexplored, and high-difficulty RL data are scarce.
Opening excerpt from the authors’ abstract. source
Read the paper
Topics
reinforcement learning large language model chain-of-thought language model long context reasoning