Opening the paper…
Figure from the original question paper Figure from the original question paper In the context of a multi-armed bandit problem with stationary reward distributions, consider the following:\ **Assertion:** UCB minimizes the regret better than the softmax approach.\ **Reason:** Softmax approach assigns a low probability of picking a sub-optimal arm that has a very low expected reward.