Reinforcement Learning End Term: 3 September 2023, Set QPE1-S1 (May 2023 term)
The IIT Madras BS Reinforcement Learning (Reinforcement Learning) End Term paper sat on 3 Sept 2023, in the May 2023 term, set QPE1-S1: 17 questions for 50 marks in 180 minutes. Every question is below with its answer. Take it as a timed mock test to be marked, or read it through first.
- 17
- 50
- 180 min
- 8
- 2
- 7
Show answer
Correct answer
Question 2
Show answer
Correct answer
Question 3
Select the most appropriate statement concerning the behaviour policy in Q-learning.
Show answer
Correct answer
Question 4
Show answer
Correct answer
Question 5
Given a problem with a well defined hierarchy, what ordering would you expect on the total expected reward for a hierarchically optimal policy (H), a recursively optimal policy (R) and a flat optimal policy (F)?
Show answer
Correct answer
Question 6
Which of the following corresponds to an update for the actor in the case of one-step actor-critic method?
Show answer
Correct answer
Question 7
Show answer
Correct answers
Question 8
Select all true statements.
Policy gradient methods use a parameterized policy that can select actions without consulting a value function.
Policy gradient methods must use a value function to learn the policy parameters.
According to the policy gradient theorem, computing the gradient of the performance requires the computation of the gradient of the state distribution μ(s).
In REINFORCE with baseline, the baseline cannot be a function of the actions.
Show answer
Correct answers
Policy gradient methods use a parameterized policy that can select actions without consulting a value function.
In REINFORCE with baseline, the baseline cannot be a function of the actions.
Question 9
Show answer
Correct answer: 0.315 (accepted within ±0.015)
Question 10
Show answer
Correct answer: 3
Question 11
Show answer
Correct answer: 0.51
Question 12
Show answer
Correct answer: 0.9
Question 13
Which of the following is the TD error used in the TD(0) algorithm?
Show answer
Correct answer
Question 14
Based on the above data, answer the given subquestions.
Find the TD error for this transition.
Show answer
Correct answer: 10
Question 15
Based on the above data, answer the given subquestions.
Show answer
Correct answer: 1
Question 16
Based on the above data, answer the given subquestions.
Show answer
Correct answer: 0.725 (accepted within ±0.015)
Question 17
Based on the above data, answer the given subquestions.
If the eligibility traces are replacing in nature, which state would have highest eligibility trace at the end of a trajectory, and what will be the value of the eligibility trace of the corresponding state?
Show answer
Correct answer
