Opening the paper…
Large Language Models, End Term
In the standard Transformer Decoder, the Multi-Head Attention layer is "Masked". What is the specific purpose of this mask during training?
In the standard Transformer Decoder, the Multi-Head Attention layer is "Masked". What is the specific purpose of this mask during training? Question text from the original paper, with its maths as pictures Question text from the original paper, with its maths as pictures