Question 7
Which of the following are components found within a standard Transformer Encoder layer?
Multi-Head Self-Attention mechanism
Position-wise Feed-Forward Networks
Cross-Attention mechanism (Encoder-Decoder attention)
Masked Multi-Head Self-Attention