A Principled Path to Fitted Distributional Evaluation
NeurIPSSpotlight2025
TL;DR
In reinforcement learning, distributional off-policy evaluation (OPE) focuses on estimating the return distribution of a target policy using offline data collected under a different policy…
Opening excerpt from the authors’ abstract. source
Read the paper
Topics
reinforcement learning evaluation
← All NeurIPS 2025 Spotlight papers · Browse the whole archive