Question 3
DDPG uses an experience replay buffer. What is the direct consequence of using an experience replay buffer
The policy gradient estimate becomes unbiased
The algorithm becomes on-policy
Temporal correlations between consecutive samples are broken
The critic no longer requires a target network