Overthinking Reduction with Decoupled Rewards and Curriculum Data Scheduling
ICLROral2026
TL;DR
theoretical unveil the underlying limitations of length reward and propose D$^2$yOR to achieve supreme efficiency without performance degradation…
Opening excerpt from the authors’ abstract. source