Quiz Space

Reinforcement Learning · End Term · 13 Apr 2025 · January 2025 term · Set 1

Question 3: Choose the correct statement in context of multi armed ba…

Question 3

+3 marksOne or more correct options

Choose the correct statement in context of multi armed bandits (MAB), assuming stationary and normal reward distribution:

Select all that apply.

  1. A

    If there are n arms, the lower bound of finding the optimal arm is ω(n).

  2. B

    If an agent doesn’t sufficiently pull each arm, then it can incorrectly pick asuboptimal arm as the optimal arm.

  3. C

    Exploration is very important step and an agent should keep exploringregularly, to minimize the regret.

  4. D

    The optimal arm can be determined by pulling each arm once and then itshould be pulled every time afterwards.

Show answer

Correct answers

  • A

    If there are n arms, the lower bound of finding the optimal arm is ω(n).

  • B

    If an agent doesn’t sufficiently pull each arm, then it can incorrectly pick asuboptimal arm as the optimal arm.

Question 3 of 16 in the IIT Madras BS Reinforcement Learning (Reinforcement Learning) End Term paper sat on 13 Apr 2025, in the January 2025 term (IIT M IMPROVEMENT FN EXAM QIM2 13 Apr). It carries 3 marks.

This question was also asked in

More questions from this paper

  1. Q1In which of the following scenarios is Expected SARSA a good fit?
  2. Q2Figure question
  3. Q4Which of the following methods are a form of Generalized Policy Iteration?
  4. Q5Figure question
  5. Q6Figure question
  6. Q7Figure question
  7. Q8Based on the above data, answer the given subquestions.
  8. Q9Based on the above data, answer the given subquestions.
  9. Q10Based on the above data, answer the given subquestions.
  10. Q11Based on the above data, answer the given subquestions.
  11. Q12Which of the routes illustrated on the grid is taken when the Hierarchically Optimal policy is executed?
  12. Q13Which of the routes illustrated on the grid is taken when the Flat Optimal policy is executed?
  13. Q14Calculate the number of time-steps taken to finish the episode when the re-cursively optimal policy is executed.
  14. Q15Based on the above data, answer the given subquestions.
  15. Q16Which action would you pick, after having performed this rollout?