Quiz Space

Reinforcement Learning · End Term · 13 Apr 2025 · January 2025 term · Set 1-5

Question 1: Consider the following assertion reason pair and select t…

Question 1

+2 marksOne correct option

Consider the following assertion reason pair and select the correct option:
Assertion: Reinforcement learning is a type of unsupervised learning algorithm as both don’t have correct labels.
Reason: In unsupervised learning, a reward like quantity is not maximized.

  1. A

    Assertion and Reason are both true and Reason is a correct explanation ofAssertion.

  2. B

    Assertion and Reason are both true and Reason is not a correct explanation ofAssertion.

  3. C

    Assertion is true but Reason is false.

  4. D

    Assertion is false but Reason is true.

Show answer

Correct answer

  • D

    Assertion is false but Reason is true.

Question 1 of 16 in the IIT Madras BS Reinforcement Learning (Reinforcement Learning) End Term paper sat on 13 Apr 2025, in the January 2025 term (IIT M IMPROVEMENT FN EXAM QIM2 13 Apr). It carries 2 marks.

This question was also asked in

More questions from this paper

  1. Q2Consider a reinforcement learning agent navigating a grid world environment. The agent receives rewards of +1 for reach…
  2. Q3In Q-learning, how does maximization bias affect the performance of the algorithm in complex environments?
  3. Q4In which of the following scenarios is Expected SARSA a good fit?
  4. Q5Which of the following methods are a form of Generalized Policy Iteration?
  5. Q6In the context of actor-critic methods, what is the effect of replacing the return Gt with the TD target?
  6. Q7Figure question
  7. Q8Figure question
  8. Q9Based on the above data, answer the given subquestions.
  9. Q10Based on the above data, answer the given subquestions.
  10. Q11Based on the above data, answer the given subquestions.
  11. Q12Based on the above data, answer the given subquestions.
  12. Q13Based on the above data, answer the given subquestions.
  13. Q14Which of the routes illustrated on the grid is taken when the Recursively Optimal policy is executed?
  14. Q15Which of the routes illustrated on the grid is taken when the Flat Optimal policy is executed?
  15. Q16Calculate the number of time-steps taken to finish the episode when the hierarchically optimal policy is executed.