Quiz Space

Reinforcement Learning · End Term · 10 May 2026 · January 2026 term · Set S2

Reinforcement Learning End Term 10 May 2026 — Question 21

Question 21

+2 marksNumerical answer
Show answer

Correct answer: 2.74 (accepted within ±0.01)

Question 21 of 21 in the IIT Madras BS Reinforcement Learning (Reinforcement Learning) End Term paper sat on 10 May 2026, in the January 2026 term (Reinforcement Learning 10 May 26 (Session 2)). It carries 2 marks.

More questions from this paper

  1. Q1Figure question
  2. Q2In DDPG, exploration is typically achieved by:
  3. Q3Figure question
  4. Q4Figure question
  5. Q5[Bonus] TRPO uses a backtracking line search after computing the natural gradient direction. What is the primary purpos…
  6. Q6Which of the following statements about DDPG are correct?
  7. Q7Which of the following are not valid conditions required for the DPG theorem to hold?
  8. Q8Which of the following statements are correct? Select all that apply.
  9. Q9Figure question
  10. Q10Figure question
  11. Q11Figure question
  12. Q12Aditi's unstable DDPG implementation is likely to suffer from which of the following problems? Select all that apply.
  13. Q13Figure question
  14. Q14Which of the following statements about the eligibility vector for a Bernoulli-logistic unit are correct? Select all th…
  15. Q15Figure question
  16. Q16Why does the REINFORCE gradient estimator exhibit high variance, particularly in long-horizon robotic tasks?
  17. Q17Figure question
  18. Q18Figure question
  19. Q19Which of the following is a key advantage of Intra-Option Q-Learning over SMDP Q-Learning?
  20. Q20Figure question