Diversity-Aware Policy Optimization for Large Language Model Reasoning

NeurIPSSpotlight2025

Authors
Jian Yao, Ran Cheng, Xingyu Wu, Jibin Wu, KC Tan
Venue
NeurIPS 2025
Track
Spotlight

TL;DR

The reasoning capabilities of large language models (LLMs) have advanced rapidly, particularly following the release of DeepSeek-R1, which has inspired a surge of research into data quality and reinfo…

Opening excerpt from the authors’ abstract. source

Read the paper

Topics

large language model language model optimization reasoning llm

← All NeurIPS 2025 Spotlight papers · Browse the whole archive