Question 9
Consider the following statements, all of which are regarding policy iteration run on a finite MDP: (1) Policy iteration can be used to find a deterministic optimal policy.
(2) Policy iteration is an algorithm that is exclusively used to evaluate the value function for a given policy.
(3) We can use the optimal value function output by policy iteration to find out all possible optimal policies, both deterministic and stochastic.
How many of these statements are true?