Question 16
Based on the above data, answer the given subquestions.
Calculate the number of time-steps taken to finish the episode when the hierarchically optimal policy is executed.
Based on the above data, answer the given subquestions.
Calculate the number of time-steps taken to finish the episode when the hierarchically optimal policy is executed.
Correct answer: 10
Question 16 of 16 in the IIT Madras BS Reinforcement Learning (Reinforcement Learning) End Term paper sat on 13 Apr 2025, in the January 2025 term (IIT M IMPROVEMENT FN EXAM QIM2 13 Apr). It carries 3 marks.