Question 6
In a -armed bandit setting, we maintain a running estimate of the action-value for each arm , denoted . We now consider using an upper confidence bound (UCB) style rule for arm selection at time , where is the number of times arm has been selected up to time , and is a tunable hyperparameter. Which of the following are good strategies for arm selection?




None of these