Question 1
The discount factor used during option execution
The discount factor used during option execution
In DDPG, exploration is typically achieved by:
Randomly sampling from the replay buffer with higher priority
Injecting noise into the critic network parameters
Adding time-correlated noise (e.g., Ornstein-Uhlenbeck) or Gaussian noise to the deterministic policy output
Reduces variance without introducing bias
Eliminates all bias in the gradient estimate
Makes the algorithm off-policy
Removes the need for a value function approximator
Sign in with Google — it is free — to see every question with its answer and explanation, practise it in learning mode, or take it as a timed mock test.
The IIT Madras BS Reinforcement Learning (Reinforcement Learning) End Term paper sat on 10 May 2026, in the January 2026 term, set S2: 21 questions for 30 marks in 180 minutes. The first 3 questions are below. Sign in with Google — it is free — to see the whole paper with its answers and explanations, in learning mode or as a timed mock test.
| Feature | Reinforcement Learning End Term 10 May 2026 Set S2 at a glance |
|---|---|
| Term | January 2026 term |
| Subject | Reinforcement Learning |
| Course code | BSDA5007 |
| Questions | 21 |
| Marks | 30 |
| Duration | 180 min |
| MCQ | 8 |
| MSQ | 7 |
| Numerical | 6 |
| Official paper | Reinforcement Learning 10 May 26 (Session 2) |
| Negative marking | No negative marking. |
| Updated |