Question 8
Recall standard control algorithms in RL, like Q-Learning, SARSA, Expected SARSA, etc., involve a maximisation step in constructing their target policies.
For example: In Q-Learning, the target is:
Which of the following statements are correct? Select all that apply.
(A figure from the original paper is missing from the source site.)
(A figure from the original paper is missing from the source site.)
(A figure from the original paper is missing from the source site.)