Large Language Models, End Term
Consider the following Assertion (A) and Reason (R) about the Transformer encoder. Assertion (A): In a Transformer encoder, each token can attend to every other token in the input sequence through self-attention. Reason (R): The encoder's self-attention mechanism uses a causal mask to prevent each token from attending to tokens that appear later in the sequence. Choose the correct option:
Consider the following Assertion (A) and Reason (R) about the Transformer encoder. Assertion (A): In a Transformer encoder, each token can attend to every other token in the input sequence through self-attention. Reason (R): The encoder's self-attention mechanism uses a causal mask to prevent each token from attending to tokens that appear later in the sequence. Choose the correct option: Which of the following statements regarding one-hot positional encoding and sinusoidal positional encoding is incorrect? You are designing an LLM to translate a sentence from English to Spanish with low computational complexity. The system has to consider a few possible sequences of words rather than committing to the single highest-probability word at every step. Which of the following decoding strategies is the most appropriate?