Quiz Space

Reinforcement Learning · End Term · 10 May 2026 · January 2026 term · Set S2

Question 19: Which of the following is a key advantage of Intra-Optio…

Question 19

+2 marksOne correct option

Which of the following is a key advantage of Intra-Option Q-Learning over SMDP Q-Learning?

  1. A

    Intra-Option Q-Learning converges to the globally optimal policy, whereas SMDP Q-Learning only converges to locally optimal solutions

  2. B
  3. C
  4. D

    Intra-Option Q-Learning learns only a single option, avoiding the need for a high-level policy

Show answer

Correct answer

  • B

Question 19 of 21 in the IIT Madras BS Reinforcement Learning (Reinforcement Learning) End Term paper sat on 10 May 2026, in the January 2026 term (Reinforcement Learning 10 May 26 (Session 2)). It carries 2 marks.

More questions from this paper

  1. Q1Figure question
  2. Q2In DDPG, exploration is typically achieved by:
  3. Q3Figure question
  4. Q4Figure question
  5. Q5[Bonus] TRPO uses a backtracking line search after computing the natural gradient direction. What is the primary purpos…
  6. Q6Which of the following statements about DDPG are correct?
  7. Q7Which of the following are not valid conditions required for the DPG theorem to hold?
  8. Q8Which of the following statements are correct? Select all that apply.
  9. Q9Figure question
  10. Q10Figure question
  11. Q11Figure question
  12. Q12Aditi's unstable DDPG implementation is likely to suffer from which of the following problems? Select all that apply.
  13. Q13Figure question
  14. Q14Which of the following statements about the eligibility vector for a Bernoulli-logistic unit are correct? Select all th…
  15. Q15Figure question
  16. Q16Why does the REINFORCE gradient estimator exhibit high variance, particularly in long-horizon robotic tasks?
  17. Q17Figure question
  18. Q18Figure question
  19. Q20Figure question
  20. Q21Figure question