Diversity-Aware Policy Optimization for Large Language Model Reasoning
NeurIPSSpotlight2025
TL;DR
The reasoning capabilities of large language models (LLMs) have advanced rapidly, particularly following the release of DeepSeek-R1, which has inspired a surge of research into data quality and reinfo…
Opening excerpt from the authors’ abstract. source
Read the paper
Topics
large language model language model optimization reasoning llm
← All NeurIPS 2025 Spotlight papers · Browse the whole archive