Question 15
Rajesh has created a GPT-like transformer model. However he doesn't have access to large compute infrastructure, so he has configured the original architecture in the following manner:
- He has taken a vocabulary of size 1000 words/tokens.
- Sequence length is 64.
- Embedding dimension and is 128.
- There are only 4 transformer blocks (layers).
- Each transformer block (layer) has only 4 attention heads.
- FFN hidden layer size is 256.
Based on the above data, answer the given subquestions.
What will be the shape of positional embedding?
64 × 128
64 × 4
12 × 128
512 × 768
None of these.