Pass@K Policy Optimization: Solving Harder Reinforcement Learning Problems
NeurIPSSpotlight2025
TL;DR
Reinforcement Learning algorithms commonly sample multiple ($n>1$) solution attempts for each problem and reward them independently…
Opening excerpt from the authors’ abstract. source
Read the paper
Topics
reinforcement learning optimization
← All NeurIPS 2025 Spotlight papers · Browse the whole archive