Reinforcement Learning End Term 10 May 2026 — Question 4
Show answer
Correct answer
Question 4 of 21 in the IIT Madras BS Reinforcement Learning (Reinforcement Learning) End Term paper sat on 10 May 2026, in the January 2026 term (Reinforcement Learning 10 May 26 (Session 2)). It carries 1 mark.
More questions from this paper
- Figure question
- In DDPG, exploration is typically achieved by:
- Figure question
- [Bonus] TRPO uses a backtracking line search after computing the natural gradient direction. What is the primary purpos…
- Which of the following statements about DDPG are correct?
- Which of the following are not valid conditions required for the DPG theorem to hold?
- Which of the following statements are correct? Select all that apply.
- Figure question
- Figure question
- Figure question
- Aditi's unstable DDPG implementation is likely to suffer from which of the following problems? Select all that apply.
- Figure question
- Which of the following statements about the eligibility vector for a Bernoulli-logistic unit are correct? Select all th…
- Figure question
- Why does the REINFORCE gradient estimator exhibit high variance, particularly in long-horizon robotic tasks?
- Figure question
- Figure question
- Which of the following is a key advantage of Intra-Option Q-Learning over SMDP Q-Learning?
- Figure question
- Figure question