Question 2
Which of the following is/are NOT present in a standard GPT (decoder-only) model?
Positional Embeddings
Word Embeddings
MultiHead Cross Attention Layer
MultiHead Self Attention Layer for the Encoder
MultiHead Self Attention Layer for the Decoder