Reinforcement Learning, End Term
January 2026 term, 10 May 2026, Set 1
State space, action space, reward function
Initiation set, option policy, termination function
Value function, behaviour policy, discount factor
Transition model, reward model, termination condition
Sign in to report a problem with this question.
You cannot change your answers after submitting.
The palette shows the status of every question. Pick a number to go straight to it.
Question text from the original paper, with its maths as pictures Question text from the original paper, with its maths as pictures DDPG uses an experience replay buffer. What is the direct consequence of using an experience replay buffer