Question 1
State space, action space, reward function
Initiation set, option policy, termination function
Value function, behaviour policy, discount factor
Transition model, reward model, termination condition
State space, action space, reward function
Initiation set, option policy, termination function
Value function, behaviour policy, discount factor
Transition model, reward model, termination condition
To prevent the actor from updating faster than the critic
To stabilise training by preventing the TD target from changing too rapidly
To ensure the replay buffer contains on-policy data
To match the learning rate of the actor and critic networks
DDPG uses an experience replay buffer. What is the direct consequence of using an experience replay buffer
The policy gradient estimate becomes unbiased
The algorithm becomes on-policy
Temporal correlations between consecutive samples are broken
The critic no longer requires a target network
Sign in with Google — it is free — to see every question with its answer and explanation, practise it in learning mode, or take it as a timed mock test.
The IIT Madras BS Reinforcement Learning (Reinforcement Learning) End Term paper sat on 10 May 2026, in the January 2026 term, set 1: 20 questions for 30 marks in 180 minutes. The first 3 questions are below. Sign in with Google — it is free — to see the whole paper with its answers and explanations, in learning mode or as a timed mock test.
| Feature | Reinforcement Learning End Term 10 May 2026 Set 1 at a glance |
|---|---|
| Term | January 2026 term |
| Subject | Reinforcement Learning |
| Course code | BSDA5007 |
| Questions | 20 |
| Marks | 30 |
| Duration | 180 min |
| MCQ | 7 |
| MSQ | 7 |
| Numerical | 6 |
| Official paper | Reinforcement Learning 10 May 26 (Session 2) |
| Negative marking | No negative marking. |
| Updated |