Question 22
During training, they strictly enforce teacher forcing, meaning the ground-truth previous tokens are always provided as input to predict the current token. Based on the above setup, answer the given subquestions.
Which of the following describes the fundamental training mechanism of this model?