uiz Space

May 2023 term · Reinforcement Learning · BSDA5007

Reinforcement Learning Quiz 1: 16 July 2023 (May 2023 term)

The IIT Madras BS Reinforcement Learning (Reinforcement Learning) Quiz 1 paper sat on 16 Jul 2023, in the May 2023 term: 19 questions for 50 marks in 120 minutes. Every question is below with its answer. Take it as a timed mock test to be marked, or read it through first.

Questions
19
Marks
50
Duration
120 min
MCQ
5
MSQ
3
Numerical
11

Updated

Official paper: IIT M DEGREE AN2 EXAM QPE2 16 JULY 2023 · No negative marking.

Question 1

+3 marksOne correct option
  1. A
  2. B
  3. C
  4. D
Show answer

Correct answer

  • A

Question 2

+3 marksOne correct option
  1. A
  2. B
  3. C
  4. D
Show answer

Correct answer

  • C

Question 3

+3 marksOne correct option

In the context of a multi-armed bandit problem with stationary reward distributions, consider the following:
Assertion: UCB minimizes the regret better than the softmax approach.
Reason: Softmax approach assigns a low probability of picking a sub-optimal arm that has a very low expected reward.

  1. A

    Assertion and Reason are both true and Reason is a correct explanation of the Assertion.

  2. B

    Assertion and Reason are both true and Reason is not a correct explanation of the Assertion.

  3. C

    Assertion is true and Reason is false

  4. D

    Both Assertion and Reason are false.

Show answer

Correct answer

  • B

    Assertion and Reason are both true and Reason is not a correct explanation of the Assertion.

Question 4

+2 marksOne correct option
  1. A

    TRUE

  2. B

    FALSE

Show answer

Correct answer

  • A

    TRUE

Question 5

+4 marksOne correct option
  1. A
  2. B
  3. C
  4. D
Show answer

Correct answer

  • B

Question 6

+4 marksOne or more correct options

Select all that apply.

  1. A

    1

  2. B

    2

  3. C

    3

  4. D

    4

  5. E

    5

Show answer

Correct answers

  • D

    4

  • E

    5

Question 7

+4 marksOne or more correct options

Select all that apply.

  1. A
  2. B
  3. C
  4. D
Show answer

Correct answers

  • A
  • C

Question 8

+3 marksOne or more correct options

Select all that apply.

  1. A
  2. B
  3. C
  4. D
  5. E
Show answer

Correct answers

  • A
  • D

Question 9

+4 marksNumerical answer

Consider the following statements, all of which are regarding policy iteration run on a finite MDP: (1) Policy iteration can be used to find a deterministic optimal policy.
(2) Policy iteration is an algorithm that is exclusively used to evaluate the value function for a given policy.
(3) We can use the optimal value function output by policy iteration to find out all possible optimal policies, both deterministic and stochastic.
How many of these statements are true?

Show answer

Correct answer: 2

Question 10

+4 marksNumerical answer
Show answer

Correct answer: 12.52 (accepted within ±0.01)

Question 11

+4 marksNumerical answer
Show answer

Correct answer: 0.92 (accepted within ±0.01)

Question 12

+1.5 marksNumerical answer

Based on the above data, answer the given subquestions.

Show answer

Correct answer: -6

Question 13

+1.5 marksNumerical answer

Based on the above data, answer the given subquestions.

Show answer

Correct answer: -8

Question 14

+1.5 marksNumerical answer

Based on the above data, answer the given subquestions.

Show answer

Correct answer: -6

Question 15

+1.5 marksNumerical answer

Based on the above data, answer the given subquestions.

Show answer

Correct answer: 0

Question 16

+1 markNumerical answer

Based on the above data, answer the given subquestions.

Show answer

Correct answer: -4

Question 17

+1 markNumerical answer

Based on the above data, answer the given subquestions.

How many deterministic optimal policies does this MDP have?

Show answer

Correct answer: 2

Question 18

+2 marksNumerical answer

Based on the above data, answer the given subquestions.

Show answer

Correct answer: 3

Question 19

+2 marksNumerical answer

Based on the above data, answer the given subquestions.

Show answer

Correct answer: 2.75 (accepted within ±0.01)