Quiz Space

Reinforcement Learning · Quiz 1 · 16 Jul 2023 · May 2023 term

Question 9: Consider the following statements, all of which are regar…

Question 9

+4 marksNumerical answer

Consider the following statements, all of which are regarding policy iteration run on a finite MDP: (1) Policy iteration can be used to find a deterministic optimal policy.
(2) Policy iteration is an algorithm that is exclusively used to evaluate the value function for a given policy.
(3) We can use the optimal value function output by policy iteration to find out all possible optimal policies, both deterministic and stochastic.
How many of these statements are true?

Show answer

Correct answer: 2

Question 9 of 19 in the IIT Madras BS Reinforcement Learning (Reinforcement Learning) Quiz 1 paper sat on 16 Jul 2023, in the May 2023 term (IIT M DEGREE AN2 EXAM QPE2 16 JULY 2023). It carries 4 marks.

More questions from this paper

  1. Q1Figure question
  2. Q2Figure question
  3. Q3In the context of a multi-armed bandit problem with stationary reward distributions, consider the following:\ Assertion…
  4. Q4Figure question
  5. Q5Figure question
  6. Q6Figure question
  7. Q7Figure question
  8. Q8Figure question
  9. Q10Figure question
  10. Q11Figure question
  11. Q12Based on the above data, answer the given subquestions.
  12. Q13Based on the above data, answer the given subquestions.
  13. Q14Based on the above data, answer the given subquestions.
  14. Q15Based on the above data, answer the given subquestions.
  15. Q16Based on the above data, answer the given subquestions.
  16. Q17How many deterministic optimal policies does this MDP have?
  17. Q18Based on the above data, answer the given subquestions.
  18. Q19Based on the above data, answer the given subquestions.