Question 20
Consider a 2-state MDP with states and , discount factor . There are two actions in : and . Action gives immediate reward and moves deterministically to ; action gives immediate reward and keeps the agent in . In , there is a single action that yields an immediate reward and transitions back to . Assume initial value estimates are , Based on the above data, answer the given subquestions.