Quiz Space

Reinforcement Learning End Term: 10 May 2026, Set 1 (January 2026 term)

Question 1

+1 markOne correct option
  1. A

    State space, action space, reward function

  2. B

    Initiation set, option policy, termination function

  3. C

    Value function, behaviour policy, discount factor

  4. D

    Transition model, reward model, termination condition

Question 2

+1 markOne correct option
  1. A

    To prevent the actor from updating faster than the critic

  2. B

    To stabilise training by preventing the TD target from changing too rapidly

  3. C

    To ensure the replay buffer contains on-policy data

  4. D

    To match the learning rate of the actor and critic networks

Question 3

+1 markOne correct option

DDPG uses an experience replay buffer. What is the direct consequence of using an experience replay buffer

  1. A

    The policy gradient estimate becomes unbiased

  2. B

    The algorithm becomes on-policy

  3. C

    Temporal correlations between consecutive samples are broken

  4. D

    The critic no longer requires a target network

17 more questions in this paper

Sign in with Google — it is free — to see every question with its answer and explanation, practise it in learning mode, or take it as a timed mock test.

More on the Reinforcement Learning End Term 10 May 2026 Set 1 paper

The IIT Madras BS Reinforcement Learning (Reinforcement Learning) End Term paper sat on 10 May 2026, in the January 2026 term, set 1: 20 questions for 30 marks in 180 minutes. The first 3 questions are below. Sign in with Google — it is free — to see the whole paper with its answers and explanations, in learning mode or as a timed mock test.

FeatureReinforcement Learning End Term 10 May 2026 Set 1 at a glance
TermJanuary 2026 term
SubjectReinforcement Learning
Course codeBSDA5007
Questions20
Marks30
Duration180 min
MCQ7
MSQ7
Numerical6
Official paperReinforcement Learning 10 May 26 (Session 2)
Negative markingNo negative marking.
Updated

Other sets that day

Same End Term, other subjects

More Reinforcement Learning