Quiz Space

Reinforcement Learning End Term: 13 April 2025, Set 1-5 (January 2025 term)

Question 1

+2 marksOne correct option

Consider the following assertion reason pair and select the correct option:
Assertion: Reinforcement learning is a type of unsupervised learning algorithm as both don’t have correct labels.
Reason: In unsupervised learning, a reward like quantity is not maximized.

  1. A

    Assertion and Reason are both true and Reason is a correct explanation ofAssertion.

  2. B

    Assertion and Reason are both true and Reason is not a correct explanation ofAssertion.

  3. C

    Assertion is true but Reason is false.

  4. D

    Assertion is false but Reason is true.

Also asked in End Term 13 Apr 2025

Question 2

+2 marksOne correct option

Consider a reinforcement learning agent navigating a grid world environment. The agent receives rewards of +1 for reaching the goal state and 0 otherwise. Which of the following statements accurately describes the differences between Monte Carlo and TD learning in this scenario?

  1. A

    Monte Carlo updates are unbiased estimators of the true value function, whileTD updates may introduce bias.

  2. B

    TD updates are guaranteed to converge to the optimal value function, whileMonte Carlo updates may not converge.

  3. C

    Monte Carlo updates require less memory and computational resourcescompared to TD updates.

  4. D

    TD updates are more robust to noise and stochasticity in the environmentcompared to Monte Carlo updates.

Also asked in End Term 13 Apr 2025

Question 3

+2 marksOne correct option

In Q-learning, how does maximization bias affect the performance of the algorithm in complex environments?

  1. A

    Maximization bias can lead to overestimation of action values, resulting insuboptimal policies and slower convergence to the optimal policy.

  2. B

    Maximization bias helps to accelerate learning by prioritizing actions withhigher estimated values, leading to faster convergence to the optimal policy.

  3. C

    Maximization bias reduces the exploration-exploitation trade-off, resulting inmore exploratory behavior and improved generalization to unseen states.

  4. D

    Maximization bias has minimal impact on the performance of Q-learning incomplex environments, as it tends to balance out over time through exploration.

Also asked in End Term 13 Apr 2025

13 more questions in this paper

Sign in with Google — it is free — to see every question with its answer and explanation, practise it in learning mode, or take it as a timed mock test.

More on the Reinforcement Learning End Term 13 Apr 2025 Set 1-5 paper

The IIT Madras BS Reinforcement Learning (Reinforcement Learning) End Term paper sat on 13 Apr 2025, in the January 2025 term, set 1-5: 16 questions for 43 marks in 180 minutes. The first 3 questions are below. Sign in with Google — it is free — to see the whole paper with its answers and explanations, in learning mode or as a timed mock test.

FeatureReinforcement Learning End Term 13 Apr 2025 Set 1-5 at a glance
TermJanuary 2025 term
SubjectReinforcement Learning
Course codeBSDA5007
Questions16
Marks43
Duration180 min
MCQ6
MSQ3
Numerical7
Official paperIIT M IMPROVEMENT FN EXAM QIM2 13 Apr
Negative markingNo negative marking.
Updated

Other sets that day

Same End Term, other subjects

More Reinforcement Learning