Question 6
A hidden activation h is used by two downstream branches, and both branches affect the same scalar loss L. During backpropagation, the gradient with respect to h is obtained by:
Taking the larger of the two branch gradients
Summing the gradient contributions from the two branches
Multiplying the two branch gradients
Averaging the two branch gradients regardless of the graph