Quiz Space

Reinforcement Learning End Term: 24 December 2023, Set FDB1 (September 2023 term)

Question 1

+1 markOne correct option

In the context of a multi arm bandit (MAB) problem and stationary reward distribution, consider following:
Assertion: UCB minimizes the regret better than ϵ-greedy approach.
Reason: ε-greedy approach keeps the probability of choosing a suboptimal arm constant.

  1. A

    Assertion and Reason are both true and Reason is a correct explanation of Assertion.

  2. B

    Assertion and Reason are both true and Reason is not a correct explanation of Assertion.

  3. C

    Assertion is true and Reason is false

  4. D

    Both Assertion and Reason are false.

Question 2

+1 markOne correct option

In the policy improvement step of policy iteration for a finite MDP, if the tie among actions, which have the same maximum value, is broken randomly, what would happen to the convergence of the algorithm?

  1. A

    The algorithm’s convergence is independent of how ties are broken

  2. B

    The algorithm may oscillate between multiple optimal policies and may never converge or convergence may be delayed

  3. C

    The algorithm will certainly not converge

  4. D

    None of these

Question 3

+1 markOne correct option

Real-time dynamic programming, or RTDP, is an on-policy trajectory-sampling version of which of the following algorithms?

  1. A

    Value iteration.

  2. B

    Q-learning.

  3. C

    SARSA.

  4. D

    TD.

  5. E

    None of these.

27 more questions in this paper

Sign in with Google — it is free — to see every question with its answer and explanation, practise it in learning mode, or take it as a timed mock test.

More on the Reinforcement Learning End Term 24 Dec 2023 Set FDB1 paper

The IIT Madras BS Reinforcement Learning (Reinforcement Learning) End Term paper sat on 24 Dec 2023, in the September 2023 term, set FDB1: 30 questions for 50 marks in 180 minutes. The first 3 questions are below. Sign in with Google — it is free — to see the whole paper with its answers and explanations, in learning mode or as a timed mock test.

FeatureReinforcement Learning End Term 24 Dec 2023 Set FDB1 at a glance
TermSeptember 2023 term
SubjectReinforcement Learning
Course codeBSDA5007
Questions30
Marks50
Duration180 min
MCQ18
MSQ5
Numerical7
Official paperIIT M DEGREE FN EXAM FDB1 24 Dec 2023
Negative markingNo negative marking.
Updated

Other sets that day

Same End Term, other subjects

More Reinforcement Learning