
Reinforcement Learning Quiz 1: 16 July 2023 (May 2023 term)
The IIT Madras BS Reinforcement Learning (Reinforcement Learning) Quiz 1 paper sat on 16 Jul 2023, in the May 2023 term: 19 questions for 50 marks in 120 minutes. Every question is below with its answer. Take it as a timed mock test to be marked, or read it through first.
- 19
- 50
- 120 min
- 5
- 3
- 11
Show answer
Correct answer
Question 2
Show answer
Correct answer
Question 3
In the context of a multi-armed bandit problem with stationary reward distributions, consider the following:
Assertion: UCB minimizes the regret better than the softmax approach.
Reason: Softmax approach assigns a low probability of picking a sub-optimal arm that has a very low expected reward.
Assertion and Reason are both true and Reason is a correct explanation of the Assertion.
Assertion and Reason are both true and Reason is not a correct explanation of the Assertion.
Assertion is true and Reason is false
Both Assertion and Reason are false.
Show answer
Correct answer
Assertion and Reason are both true and Reason is not a correct explanation of the Assertion.
Question 4
TRUE
FALSE
Show answer
Correct answer
TRUE
Question 5
Show answer
Correct answer
Question 6
1
2
3
4
5
Show answer
Correct answers
4
5
Question 7
Show answer
Correct answers
Question 8
Show answer
Correct answers
Question 9
Consider the following statements, all of which are regarding policy iteration run on a finite MDP: (1) Policy iteration can be used to find a deterministic optimal policy.
(2) Policy iteration is an algorithm that is exclusively used to evaluate the value function for a given policy.
(3) We can use the optimal value function output by policy iteration to find out all possible optimal policies, both deterministic and stochastic.
How many of these statements are true?
Show answer
Correct answer: 2
Question 10
Show answer
Correct answer: 12.52 (accepted within ±0.01)
Question 11
Show answer
Correct answer: 0.92 (accepted within ±0.01)
Question 12
Based on the above data, answer the given subquestions.
Show answer
Correct answer: -6
Question 13
Based on the above data, answer the given subquestions.
Show answer
Correct answer: -8
Question 14
Based on the above data, answer the given subquestions.
Show answer
Correct answer: -6
Question 15
Based on the above data, answer the given subquestions.
Show answer
Correct answer: 0
Question 16
Based on the above data, answer the given subquestions.
Show answer
Correct answer: -4
Question 17
Based on the above data, answer the given subquestions.
How many deterministic optimal policies does this MDP have?
Show answer
Correct answer: 2
Question 18
Based on the above data, answer the given subquestions.
Show answer
Correct answer: 3
Question 19
Based on the above data, answer the given subquestions.
Show answer
Correct answer: 2.75 (accepted within ±0.01)