Question 10
A transformer model has the following embedding layer: Vocabulary size = 50,000• Embedding dimension = 768• The embeddings are currently stored in float32 format.
You decide to optimize memory as follows: Quantize 70% of the embedding vectors to int81. Keep the remaining 30% in float32 (for high-frequency tokens)2. What is the overall percentage reduction in memory for the embedding matrix?
(Round to the nearest integer)
52%
60%
53%
65%