Precise Asymptotics and Refined Regret of Variance-Aware UCB
NeurIPSSpotlight2025
TL;DR
In this paper, we study the behavior of the Upper Confidence Bound-Variance (UCB-V) algorithm for the Multi-Armed Bandit (MAB) problems, a variant of the canonical Upper Confidence Bound (UCB) algorit…
Opening excerpt from the authors’ abstract. source
Read the paper
Topics
bandit regret
← All NeurIPS 2025 Spotlight papers · Browse the whole archive