Question 23
During training, they strictly enforce teacher forcing, meaning the ground-truth previous tokens are always provided as input to predict the current token. Based on the above setup, answer the given subquestions.
It dictates that the loss is only calculated for tokens that were incorrectly predicted.