A Unifying View of Linear Function Approximation in Off-Policy RL Through Matrix Splitting and Preconditioning

NeurIPSSpotlight2025

Authors
Zechen Wu, Amy Greenwald, Ronald Parr
Venue
NeurIPS 2025
Track
Spotlight

TL;DR

In off-policy policy evaluation (OPE) tasks within reinforcement learning, Temporal Difference Learning(TD) and Fitted Q-Iteration (FQI) have traditionally been viewed as differing in the number of up…

Opening excerpt from the authors’ abstract. source

Read the paper

Topics

reinforcement learning evaluation

← All NeurIPS 2025 Spotlight papers · Browse the whole archive