Pass@K Policy Optimization: Solving Harder Reinforcement Learning Problems

NeurIPSSpotlight2025

Authors
Christian Walder, Deep Tejas Karkhanis
Venue
NeurIPS 2025
Track
Spotlight

TL;DR

Reinforcement Learning algorithms commonly sample multiple ($n>1$) solution attempts for each problem and reward them independently…

Opening excerpt from the authors’ abstract. source

Read the paper

Topics

reinforcement learning optimization

← All NeurIPS 2025 Spotlight papers · Browse the whole archive