Does Stochastic Gradient really succeed for bandits?
NeurIPSOral2025
TL;DR
We propose a novel regret analysis of a simple policy gradient algorithm for bandits, characterizing regret regimes depending on its learning rate.
Opening excerpt from the authors’ abstract. source
Read the paper
Topics
bandit regret