A Principled Path to Fitted Distributional Evaluation

NeurIPSSpotlight2025

Authors
Sungee Hong, Jiayi Wang, Zhengling Qi, Raymond K. W. Wong
Venue
NeurIPS 2025
Track
Spotlight

TL;DR

In reinforcement learning, distributional off-policy evaluation (OPE) focuses on estimating the return distribution of a target policy using offline data collected under a different policy…

Opening excerpt from the authors’ abstract. source

Read the paper

Topics

reinforcement learning evaluation

← All NeurIPS 2025 Spotlight papers · Browse the whole archive