Extended UCB Policies for Multi-armed Bandit Problems
摘要
The multi-armed bandit (MAB) problems are widely studied in fields of operations research, stochastic optimization, and reinforcement learning. In this paper, we consider the classical MAB model with heavy-tailed reward distributions and introduce the extended robust UCB policy, which is an extension of the results of Bubeck et al. (IEEE Trans Inf Theory 59:7711–7717, 2013) and Lattimore (Adv Neural Inf Process Syst, 30:1–10, 2017)that are further based on the pioneering idea of UCB policies e.g. (Auer et al. in Mach Learn, 47:235–256, 2002). The previous UCB policies require some strict conditions on reward distributions, which can be difficult to guarantee in practical scenarios. Our extended robust UCB generalizes Lattimore’s seminary work (for moments of orders