Quiz Space

Reinforcement Learning Quiz 2: 23 November 2025 (September 2025 term)

Question 1

+2 marksOne correct option

Consider the following assertion reason pair:
Assertion: In Monte Carlo (MC), the value function converges to the certainty equivalence estimate.
Reason: Monte Carlo methods use complete trajectories to fit the value function as closely as possible to the sampled returns.

  1. A

    Both Assertion and Reason are correct, and Reason is the correct explanation.

  2. B

    Assertion is correct, Reason is incorrect

  3. C

    Assertion is incorrect, Reason is correct

  4. D

    Both Assertion and Reason are correct, but Reason is not the correct explanation.

Question 2

+3 marksOne or more correct options

Which of the following is the correct way to represent policy for policy search methods? Assume represents a real-valued parameter.

Select all that apply.

  1. A
  2. B
  3. C
  4. D
  5. E

Question 3

+3 marksOne or more correct options

Recall the incremental update rule for REINFORCE:

Consider the following binary-bandit problem:

Which of the following expressions are equivalent to
?

Select all that apply.

  1. A
  2. B
  3. C
  4. D

14 more questions in this paper

Sign in with Google — it is free — to see every question with its answer and explanation, practise it in learning mode, or take it as a timed mock test.

More on the Reinforcement Learning Quiz 2 23 Nov 2025 paper

The IIT Madras BS Reinforcement Learning (Reinforcement Learning) Quiz 2 paper sat on 23 Nov 2025, in the September 2025 term: 17 questions for 50 marks in 120 minutes. The first 3 questions are below. Sign in with Google — it is free — to see the whole paper with its answers and explanations, in learning mode or as a timed mock test.

FeatureReinforcement Learning Quiz 2 23 Nov 2025 at a glance
TermSeptember 2025 term
SubjectReinforcement Learning
Course codeBSDA5007
Questions17
Marks50
Duration120 min
MCQ8
MSQ5
Numerical4
Official paperIIT M DEGREE AN EXAM QDB2 23 Nov 2025 NEW
Negative markingNo negative marking.
Updated

Same Quiz 2, other subjects

More Reinforcement Learning