Quiz Space

Reinforcement Learning Quiz 1: 15 March 2026 (January 2026 term)

Question 1

+2 marksOne correct option

Which of the following is/are correct and valid reasons to consider sampling actions from a softmax distribution instead of using an greedy approach? 1.Under softmax exploration, the probability of selecting an action increases with its estimated action value, which reduces unnecessary exploration of clearly inferior actions. 2.Unlike the -greedy method, softmax exploration does not require careful, gradual decay of the exploration parameter and still yields asymptotically correct behaviour even if the temperature is reduced sharply. 3.It enables more fine-grained discrimination among actions whose estimated Q-values are close to the maximum, allowing more nuanced preference for slightly better actions. Which of the above statements is/are correct?

  1. A

    1, 2, 3

  2. B

    only 3

  3. C

    1, 2

  4. D

    1, 3

  5. E

    3, 2

  6. F

    Only 2

Question 2

+2 marksOne correct option

Consider a discounted return:

in an infinite-horizon MDP with bounded rewards and discount factor 1.For fixed rewards, the contribution of to decays geometrically as . 2.If and rewards are uniformly bounded, the infinite sum is always finite. 3.For , the discounted return is bounded above in magnitude by

. Here, Rmax be the maximum reward for any transition. Which of the above statements is/are correct?

Consider a discounted return:
Consider a discounted return:
  1. A

    1, 2, 3

  2. B

    1, 3

  3. C

    2, 3

  4. D

    only 1

  5. E

    only 3

Question 3

+1 markOne correct option

Consider the following assertion and reason pair and select the correct option: Assertion:In the UCB algorithm for multi-armed bandits, replacing the upper confidence bound with a lower confidence bound (and greedily selecting actions based on that lower bound) would still promote effective exploration or lead to optimal reward maximisation. Reason:Even when lower confidence bounds are used instead of upper bounds, if the algorithm greedily selects arms based on these lower bounds, it would still encourage exploration and ultimately achieve optimal reward maximisation.

  1. A

    Both Assertion and Reason are true, and Reason is the correct explanation of Assertion.

  2. B

    Both Assertion and Reason are true, but Reason is NOT the correct explanation of Assertion.

  3. C

    Assertion is true, Reason is false

  4. D

    Assertion is false, Reason is false

19 more questions in this paper

Sign in with Google — it is free — to see every question with its answer and explanation, practise it in learning mode, or take it as a timed mock test.

More on the Reinforcement Learning Quiz 1 15 Mar 2026 paper

The IIT Madras BS Reinforcement Learning (Reinforcement Learning) Quiz 1 paper sat on 15 Mar 2026, in the January 2026 term: 22 questions for 34 marks in 120 minutes. The first 3 questions are below. Sign in with Google — it is free — to see the whole paper with its answers and explanations, in learning mode or as a timed mock test.

FeatureReinforcement Learning Quiz 1 15 Mar 2026 at a glance
TermJanuary 2026 term
SubjectReinforcement Learning
Course codeBSDA5007
Questions22
Marks34
Duration120 min
MCQ11
MSQ1
Written10
Official paperReinforcement Learning 15 Mar 26
Negative markingNo negative marking.
Updated

Same Quiz 1, other subjects

More Reinforcement Learning