Question 3
Reduces variance without introducing bias
Eliminates all bias in the gradient estimate
Makes the algorithm off-policy
Removes the need for a value function approximator
Reduces variance without introducing bias
Eliminates all bias in the gradient estimate
Makes the algorithm off-policy
Removes the need for a value function approximator
Correct answer
Reduces variance without introducing bias
Question 3 of 21 in the IIT Madras BS Reinforcement Learning (Reinforcement Learning) End Term paper sat on 10 May 2026, in the January 2026 term (Reinforcement Learning 10 May 26 (Session 2)). It carries 1 mark.