Question 1
In the context of a multi arm bandit (MAB) problem and stationary reward distribution, consider following:
Assertion: UCB minimizes the regret better than ϵ-greedy approach.
Reason: ε-greedy approach keeps the probability of choosing a suboptimal arm constant.
Assertion and Reason are both true and Reason is a correct explanation of Assertion.
Assertion and Reason are both true and Reason is not a correct explanation of Assertion.
Assertion is true and Reason is false
Both Assertion and Reason are false.