Question 10
Based on the above data, answer the given subquestions.
Which of the following best explains why a contextual bandit approach is preferred over a standard multi-armed bandit in this scenario?
Because the optimal waiting time is the same for all VMs, regardless of their context.
Because contextual bandits can adaptively select actions based on the specific features of each VM failure event, maximising expected reward across diverse situations.
Because multi-armed bandits are unable to balance exploration and exploitation.
Because contextual bandits always guarantee zero regret after every round.