Reinforcement Learning End Term 10 May 2026 — Question 21
Show answer
Correct answer: 2.74 (accepted within ±0.01)
Question 21 of 21 in the IIT Madras BS Reinforcement Learning (Reinforcement Learning) End Term paper sat on 10 May 2026, in the January 2026 term (Reinforcement Learning 10 May 26 (Session 2)). It carries 2 marks.
More questions from this paper
- Figure question
- In DDPG, exploration is typically achieved by:
- Figure question
- Figure question
- [Bonus] TRPO uses a backtracking line search after computing the natural gradient direction. What is the primary purpos…
- Which of the following statements about DDPG are correct?
- Which of the following are not valid conditions required for the DPG theorem to hold?
- Which of the following statements are correct? Select all that apply.
- Figure question
- Figure question
- Figure question
- Aditi's unstable DDPG implementation is likely to suffer from which of the following problems? Select all that apply.
- Figure question
- Which of the following statements about the eligibility vector for a Bernoulli-logistic unit are correct? Select all th…
- Figure question
- Why does the REINFORCE gradient estimator exhibit high variance, particularly in long-horizon robotic tasks?
- Figure question
- Figure question
- Which of the following is a key advantage of Intra-Option Q-Learning over SMDP Q-Learning?
- Figure question