Opening the paper…
Figure from the original question paper Consider following assertion reason pair:\ **Assertion:**\ In the UCB algorithm for multi-armed bandits, replacing the upper confidence bound with a lower confidence bound (and greedily selecting actions based on that lower bound)\ would still promote effective exploration or lead to optimal reward maximisation.\ **Reason:**\ Even when lower confidence bounds are used instead of upper bounds, if the algorithm greedily selects arms based on these lower bounds, it would still encourage exploration and ultimately achieve optimal reward maximisation. Figure from the original question paper