Opening the paper…
Reinforcement Learning, Quiz 2
For on policy SARSA update rule for Q(s, a), s, a, s′, a′ are needed. When is a′ computed/sampled?
For on policy SARSA update rule for *Q(s, a), s, a, s′, a′* are needed. When is *a′* computed/sampled? Figure from the original question paper Figure from the original question paper