Question 41
In vision transformers, if treat each pixel as a token for an image of size 64 × 64 × 3, the number of parameters required for the self-attention calculation are ________blank(a)________ . What will be the number of parameters if consider patches of size 8 × 8 × 3 ________blank(b)________. Based on the above data, answer the given subquestions.
Enter the correct answer for blank(a):