Reinforcement Learning End Term 13 Apr 2025 — Question 8
Show answer
Correct answer: -7
Question 8 of 16 in the IIT Madras BS Reinforcement Learning (Reinforcement Learning) End Term paper sat on 13 Apr 2025, in the January 2025 term (IIT M IMPROVEMENT FN EXAM QIM2 13 Apr). It carries 3 marks.
More questions from this paper
- Consider the following assertion reason pair and select the correct option:\ Assertion: Reinforcement learning is a typ…
- Consider a reinforcement learning agent navigating a grid world environment. The agent receives rewards of +1 for reach…
- In Q-learning, how does maximization bias affect the performance of the algorithm in complex environments?
- In which of the following scenarios is Expected SARSA a good fit?
- Which of the following methods are a form of Generalized Policy Iteration?
- In the context of actor-critic methods, what is the effect of replacing the return Gt with the TD target?
- Figure question
- Based on the above data, answer the given subquestions.
- Based on the above data, answer the given subquestions.
- Based on the above data, answer the given subquestions.
- Based on the above data, answer the given subquestions.
- Based on the above data, answer the given subquestions.
- Which of the routes illustrated on the grid is taken when the Recursively Optimal policy is executed?
- Which of the routes illustrated on the grid is taken when the Flat Optimal policy is executed?
- Calculate the number of time-steps taken to finish the episode when the hierarchically optimal policy is executed.