Quiz Space

Reinforcement Learning Quiz 1: 26 October 2025 (September 2025 term)

Question 1

+3 marksOne correct option
  1. A
  2. B
  3. C
  4. D

Question 2

+3 marksOne correct option

Consider following assertion reason pair:
Assertion:
In the UCB algorithm for multi-armed bandits, replacing the upper confidence bound with a lower confidence bound (and greedily selecting actions based on that lower bound)
would still promote effective exploration or lead to optimal reward maximisation.
Reason:
Even when lower confidence bounds are used instead of upper bounds, if the algorithm greedily selects arms based on these lower bounds, it would still encourage exploration and ultimately achieve optimal reward maximisation.

  1. A

    Both Assertion and Reason are true, and Reason is the correct explanation of Assertion.

  2. B

    Both Assertion and Reason are true, but Reason is NOT the correct explanation of Assertion.

  3. C

    Assertion is true, Reason is false

  4. D

    Assertion is false, Reason is false

Question 3

+5 marksOne correct option
  1. A

    Policy Iteration

  2. B

    Value Iteration

  3. C

    Both require equal updates

  4. D

    It cannot be determined

13 more questions in this paper

Sign in with Google — it is free — to see every question with its answer and explanation, practise it in learning mode, or take it as a timed mock test.

More on the Reinforcement Learning Quiz 1 26 Oct 2025 paper

The IIT Madras BS Reinforcement Learning (Reinforcement Learning) Quiz 1 paper sat on 26 Oct 2025, in the September 2025 term: 16 questions for 50 marks in 120 minutes. The first 3 questions are below. Sign in with Google — it is free — to see the whole paper with its answers and explanations, in learning mode or as a timed mock test.

FeatureReinforcement Learning Quiz 1 26 Oct 2025 at a glance
TermSeptember 2025 term
SubjectReinforcement Learning
Course codeBSDA5007
Questions16
Marks50
Duration120 min
MCQ11
MSQ2
Numerical3
Official paperIIT M DEGREE AN EXAM QDB2 26 Oct 2025
Negative markingNo negative marking.
Updated

Same Quiz 1, other subjects

More Reinforcement Learning