Question 1
Based on the above data, answer the given subquestions.

The IIT Madras BS Reinforcement Learning (Reinforcement Learning) Quiz 1 paper sat on 7 Jul 2024, in the May 2024 term: 17 questions for 40 marks in 120 minutes. Every question is below with its answer. Take it as a timed mock test to be marked, or read it through first.
Based on the above data, answer the given subquestions.
Correct answer: 0.1 (accepted within ±0.005)
Based on the above data, answer the given subquestions.
Correct answer: 0.8 (accepted within ±0.005)
Based on the above data, answer the given subquestions.
Arm A1 is the optimal arm.
Arm A2 is the optimal arm.
Arm A3 is the optimal arm.
An optimal arm can not be determined.
There is a tie for the optimal arm.
Correct answer
Arm A3 is the optimal arm.
Based on the above data, answer the given subquestions.
Arm A1 has to provide a reward of 4 or more.
Arm A2 has to provide a reward strictly more than 6.
Arm A3 has to provide a reward strictly less than 0.
It can not be determined.
None of these.
Correct answer
Arm A2 has to provide a reward strictly more than 6.
Correct answer: 3.125 (accepted within ±0.005)
Correct answer: 3.25 (accepted within ±0.005)
Correct answer: 1.4 (accepted within ±0.005)
Correct answer: 0.665 (accepted within ±0.005)
Correct answer: 2.825 (accepted within ±0.005)
Which of the following is the correct Bellman equation for stochastic transitions, stochastic policy and stochastic rewards? The symbols have the usual meaning.
Correct answer
Based on the above data, answer the given subquestions.
Correct answer: -4
Based on the above data, answer the given subquestions.
Correct answer: -7
Based on the above data, answer the given subquestions.
1 and 2.
1 and 3.
2 and 3.
2 and 5.
13 and 9
4 and 8.
None of these
Correct answers
2 and 5.
13 and 9
4 and 8.
Correct answer: 10
Correct answer: 2
Correct answer: 16.08 (accepted within ±0.05)
Correct answer: 12.24 (accepted within ±0.01)