Question 4
In Q-Learning, the update rule for the action-value function is based on bootstrapping from the current estimate. Which of the following correctly represents this update?
In Q-Learning, the update rule for the action-value function is based on bootstrapping from the current estimate. Which of the following correctly represents this update?
Correct answers
Question 4 of 17 in the IIT Madras BS Reinforcement Learning (Reinforcement Learning) Quiz 2 paper sat on 23 Nov 2025, in the September 2025 term (IIT M DEGREE AN EXAM QDB2 23 Nov 2025 NEW). It carries 2 marks.