Question 18
What is the key advantage of Intra-Option Q-Learning (Algorithm B) over SMDP Q-Learning (Algorithm A)?
It converges to a globally optimal policy, whereas SMDP Q-learning only converges locally
It eliminates the need for a high-level policy by learning all options simultaneously