Quiz Space

Reinforcement Learning · Quiz 1 · 15 Mar 2026 · January 2026 term

Question 3: Consider the following assertion and reason pair and sele…

Question 3

+1 markOne correct option

Consider the following assertion and reason pair and select the correct option: Assertion:In the UCB algorithm for multi-armed bandits, replacing the upper confidence bound with a lower confidence bound (and greedily selecting actions based on that lower bound) would still promote effective exploration or lead to optimal reward maximisation. Reason:Even when lower confidence bounds are used instead of upper bounds, if the algorithm greedily selects arms based on these lower bounds, it would still encourage exploration and ultimately achieve optimal reward maximisation.

  1. A

    Both Assertion and Reason are true, and Reason is the correct explanation of Assertion.

  2. B

    Both Assertion and Reason are true, but Reason is NOT the correct explanation of Assertion.

  3. C

    Assertion is true, Reason is false

  4. D

    Assertion is false, Reason is false

Show answer

Correct answer

  • D

    Assertion is false, Reason is false

Question 3 of 22 in the IIT Madras BS Reinforcement Learning (Reinforcement Learning) Quiz 1 paper sat on 15 Mar 2026, in the January 2026 term (Reinforcement Learning 15 Mar 26). It carries 1 mark.

More questions from this paper

  1. Q1Which of the following is/are correct and valid reasons to consider sampling actions from a softmax distribution instea…
  2. Q2Consider a discounted return: in an infinite-horizon MDP with bounded rewards and discount factor 1.For fixed rewards, …
  3. Q4In the context of MDPs, what does it mean for a policy to be greedy with respect to an action- value function ?
  4. Q5In policy iteration for finite MDPs, how are the Bellman equations used during the policy evaluation step?
  5. Q6In a -armed bandit setting, we maintain a running estimate of the action-value for each arm , denoted . We now consider…
  6. Q7Which of the following equations best represents the Bellman optimality equation for the optimal state-value function
  7. Q8The sequence of rewards for a continuing task with is given below: Find the return . Your answer should have exactly tw…
  8. Q9You have a 3-armed bandit where each arm, when active, yields a Bernoulli reward with success probabilities , , and . H…
  9. Q10Assume the initial value estimates are , , and . After performing one synchronous policy evaluation update (under the f…
  10. Q11Starting from , , and . perform two successive synchronous policy evaluation sweeps (under the fixed policy that follow…
  11. Q12If you perform a single value iteration Bellman optimality update for , what is the new value ?
  12. Q13Suppose after some value iterations your current estimates are , . If you perform one value iteration update for , what…
  13. Q14Consider a faulty -armed bandit with -noise activation and base arm success probabilities . If you always choose the ar…
  14. Q15Consider a faulty -armed bandit with -noise activation and distinct base success probabilities . Suppose the agent foll…
  15. Q16In a faulty -armed bandit with -noise activation, suppose you mistakenly model the environment as a standard bandit wit…
  16. Q17Which of the following best explains why a contextual bandit approach is preferred over a standard multi-armed bandit i…
  17. Q18After many rounds of user interaction, what is the main optimality objective of the contextual bandit algorithm in this…
  18. Q19Consider a Markov Decision Process (MDP) with three states , , and , where is a terminal state and . Each non-terminal …
  19. Q20Consider a 2-state MDP with states and , discount factor . There are two actions in : and . Action gives immediate rewa…
  20. Q21Consider an N-armed bandit problem in which each arm , for , produces a reward of 1 with probability and otherwise, whe…
  21. Q22Consider a large-scale news recommendation platform that uses a contextual bandit approach to decide which headline to …