Does Stochastic Gradient really succeed for bandits?

NeurIPSOral2025

Authors
Dorian Baudry, Emmeran Johnson, Simon Vary, Ciara Pike-Burke, Patrick Rebeschini
Affiliation
INRIA
Venue
NeurIPS 2025
Track
Oral

TL;DR

We propose a novel regret analysis of a simple policy gradient algorithm for bandits, characterizing regret regimes depending on its learning rate.

Opening excerpt from the authors’ abstract. source

Read the paper

Topics

bandit regret

← All NeurIPS 2025 Oral papers · Browse the whole archive