Mean Flow Policy with Instantaneous Velocity Constraint for One-step Action Generation
ICLROral2026
TL;DR
We introduce the mean velocity policy, a new RL policy that, along with a novel instantaneous velocity constraint, achieves state-of-the-art performance and the fastest training and inference speed.
Opening excerpt from the authors’ abstract. source