Quiz Space

Reinforcement Learning · Quiz 1 · 15 Mar 2026 · January 2026 term

Question 9: You have a 3-armed bandit where each arm, when active, yi…

Question 9

+2 marksWritten answer

You have a 3-armed bandit where each arm, when active, yields a Bernoulli reward with success probabilities , , and . However, arm 1 is active in any given round with probability , and otherwise it yields a reward of 0 regardless of its Bernoulli outcome. Arms 2 and 3 are always active. What is the expected reward from pulling arm 1 in a single round?

Show answer

A written answer, not marked automatically.

Question 9 of 22 in the IIT Madras BS Reinforcement Learning (Reinforcement Learning) Quiz 1 paper sat on 15 Mar 2026, in the January 2026 term (Reinforcement Learning 15 Mar 26). It carries 2 marks.

More questions from this paper

  1. Q1Which of the following is/are correct and valid reasons to consider sampling actions from a softmax distribution instea…
  2. Q2Consider a discounted return: in an infinite-horizon MDP with bounded rewards and discount factor 1.For fixed rewards, …
  3. Q3Consider the following assertion and reason pair and select the correct option: Assertion:In the UCB algorithm for mult…
  4. Q4In the context of MDPs, what does it mean for a policy to be greedy with respect to an action- value function ?
  5. Q5In policy iteration for finite MDPs, how are the Bellman equations used during the policy evaluation step?
  6. Q6In a -armed bandit setting, we maintain a running estimate of the action-value for each arm , denoted . We now consider…
  7. Q7Which of the following equations best represents the Bellman optimality equation for the optimal state-value function
  8. Q8The sequence of rewards for a continuing task with is given below: Find the return . Your answer should have exactly tw…
  9. Q10Assume the initial value estimates are , , and . After performing one synchronous policy evaluation update (under the f…
  10. Q11Starting from , , and . perform two successive synchronous policy evaluation sweeps (under the fixed policy that follow…
  11. Q12If you perform a single value iteration Bellman optimality update for , what is the new value ?
  12. Q13Suppose after some value iterations your current estimates are , . If you perform one value iteration update for , what…
  13. Q14Consider a faulty -armed bandit with -noise activation and base arm success probabilities . If you always choose the ar…
  14. Q15Consider a faulty -armed bandit with -noise activation and distinct base success probabilities . Suppose the agent foll…
  15. Q16In a faulty -armed bandit with -noise activation, suppose you mistakenly model the environment as a standard bandit wit…
  16. Q17Which of the following best explains why a contextual bandit approach is preferred over a standard multi-armed bandit i…
  17. Q18After many rounds of user interaction, what is the main optimality objective of the contextual bandit algorithm in this…
  18. Q19Consider a Markov Decision Process (MDP) with three states , , and , where is a terminal state and . Each non-terminal …
  19. Q20Consider a 2-state MDP with states and , discount factor . There are two actions in : and . Action gives immediate rewa…
  20. Q21Consider an N-armed bandit problem in which each arm , for , produces a reward of 1 with probability and otherwise, whe…
  21. Q22Consider a large-scale news recommendation platform that uses a contextual bandit approach to decide which headline to …