Regret Bounds for Adversarial Contextual Bandits with General Function Approximation and Delayed Feedback
NeurIPSSpotlight2025
TL;DR
We present regret minimization algorithms for the contextual multi-armed bandit (CMAB) problem over $K$ actions in the presence of delayed feedback, a scenario where loss observations arrive with dela…
Opening excerpt from the authors’ abstract. source
Read the paper
Topics
adversarial bandit regret
← All NeurIPS 2025 Spotlight papers · Browse the whole archive