Question 2
Question 3
In the context of a multi-armed bandit problem with stationary reward distributions, consider the following:
Assertion: UCB minimizes the regret better than the softmax approach.
Reason: Softmax approach assigns a low probability of picking a sub-optimal arm that has a very low expected reward.
Assertion and Reason are both true and Reason is a correct explanation of the Assertion.
Assertion and Reason are both true and Reason is not a correct explanation of the Assertion.
Assertion is true and Reason is false
Both Assertion and Reason are false.
16 more questions in this paper
Sign in with Google — it is free — to see every question with its answer and explanation, practise it in learning mode, or take it as a timed mock test.
More on the Reinforcement Learning Quiz 1 16 Jul 2023 paper
The IIT Madras BS Reinforcement Learning (Reinforcement Learning) Quiz 1 paper sat on 16 Jul 2023, in the May 2023 term: 19 questions for 50 marks in 120 minutes. The first 3 questions are below. Sign in with Google — it is free — to see the whole paper with its answers and explanations, in learning mode or as a timed mock test.
| Feature | Reinforcement Learning Quiz 1 16 Jul 2023 at a glance |
|---|---|
| Term | May 2023 term |
| Subject | Reinforcement Learning |
| Course code | BSDA5007 |
| Questions | 19 |
| Marks | 50 |
| Duration | 120 min |
| MCQ | 5 |
| MSQ | 3 |
| Numerical | 11 |
| Official paper | IIT M DEGREE AN2 EXAM QPE2 16 JULY 2023 |
| Negative marking | No negative marking. |
| Updated |