Question 19
Which of the following is a key advantage of Intra-Option Q-Learning over SMDP Q-Learning?
Intra-Option Q-Learning converges to the globally optimal policy, whereas SMDP Q-Learning only converges to locally optimal solutions
Intra-Option Q-Learning learns only a single option, avoiding the need for a high-level policy