Quiz Space

Reinforcement Learning End Term: 13 April 2025, Set 1-3 (January 2025 term)

Question 1

+2 marksOne correct option

In which of the following scenarios is Expected SARSA a good fit?

  1. A

    When the environment has a high degree of stochasticity, makingbootstrapping unstable in value-based methods

  2. B

    When function approximation is necessary due to large state spaces, requiringdeep networks for learning representations

  3. C

    When learning needs to prioritize separating state value and advantagefunctions for better decision-making

  4. D

    When experience replay is essential for stable learning and sample efficiency

Also asked in End Term 13 Apr 2025, End Term 13 Apr 2025, End Term 13 Apr 2025, End Term 13 Apr 2025 and 1 more

Question 2

+3 marksOne correct option
  1. A

    SARSA updates its action-value function based on the action actually taken inthe next state, while Q-learning updates its action-value function based on the maximum action-value in the next state.

  2. B

    SARSA is guaranteed to converge to the optimal policy under certainconditions, while Q-learning may diverge or oscillate without additional modifications.

  3. C

    SARSA is more computationally efficient than Q-learning, requiring fewerupdates to converge to the optimal policy.

  4. D

    SARSA and Q-learning exhibit similar performance in terms of convergencespeed and solution quality in this scenario.

Question 3

+3 marksOne or more correct options

Choose the correct statement in context of multi armed bandits (MAB), assuming stationary and normal reward distribution:

Select all that apply.

  1. A

    If there are n arms, the lower bound of finding the optimal arm is ω(n).

  2. B

    If an agent doesn’t sufficiently pull each arm, then it can incorrectly pick asuboptimal arm as the optimal arm.

  3. C

    Exploration is very important step and an agent should keep exploringregularly, to minimize the regret.

  4. D

    The optimal arm can be determined by pulling each arm once and then itshould be pulled every time afterwards.

Also asked in End Term 13 Apr 2025, End Term 13 Apr 2025, End Term 13 Apr 2025

13 more questions in this paper

Sign in with Google — it is free — to see every question with its answer and explanation, practise it in learning mode, or take it as a timed mock test.

More on the Reinforcement Learning End Term 13 Apr 2025 Set 1-3 paper

The IIT Madras BS Reinforcement Learning (Reinforcement Learning) End Term paper sat on 13 Apr 2025, in the January 2025 term, set 1-3: 16 questions for 43 marks in 180 minutes. The first 3 questions are below. Sign in with Google — it is free — to see the whole paper with its answers and explanations, in learning mode or as a timed mock test.

FeatureReinforcement Learning End Term 13 Apr 2025 Set 1-3 at a glance
TermJanuary 2025 term
SubjectReinforcement Learning
Course codeBSDA5007
Questions16
Marks43
Duration180 min
MCQ5
MSQ4
Numerical7
Official paperIIT M IMPROVEMENT FN EXAM QIM2 13 Apr
Negative markingNo negative marking.
Updated

Other sets that day

Same End Term, other subjects

More Reinforcement Learning