Question 11
The model will completely ignore the future tokens during training.
The model will stop attending to any tokens in the sequence.
The model will incorrectly attend to future tokens, leading to information leakage and breaking the causality constraint.
The softmax function will fail to compute attention scores for any position.