Quiz Space

Large Language Models · End Term · 31 Aug 2025 · May 2025 term · Set QDB3

Question 4: Which of the following preprocessing steps are commonly u…

Question 4

+3 marksOne or more correct options

Which of the following preprocessing steps are commonly used when preparing data for transformer models like BERT?

Select all that apply.

  1. A

    Adding special tokens such as [CLS] and [SEP] to the sequence

  2. B

    Splitting tokens into subword units using a model-specific tokenizer

  3. C

    Removing all punctuation marks to ensure cleaner embeddings

  4. D

    Adding positional information to each token embedding

Show answer

Correct answers

  • A

    Adding special tokens such as [CLS] and [SEP] to the sequence

  • B

    Splitting tokens into subword units using a model-specific tokenizer

  • D

    Adding positional information to each token embedding

Question 4 of 17 in the IIT Madras BS Large Language Models (LLM) End Term paper sat on 31 Aug 2025, in the May 2025 term (IIT M DEGREE AN EXAM QDB3 31 Aug 2025). It carries 3 marks.

More questions from this paper

  1. Q1Given the input string: sunshine And the following vocabulary of subword tokens with their corresponding log-probabilit…
  2. Q2A GPT model is trained using causal language modeling. During training, for a sequence of T = 4 tokens, which of the fo…
  3. Q3Figure question
  4. Q5Which modifications are used in transformers to handle longer sequences efficiently?
  5. Q6In Transformer models that use relative position embeddings (such as in Transformer-XL or T5), clipping is often applie…
  6. Q7Figure question
  7. Q8For the input “you enjoy tea often”, compute the final representation of the word “enjoy” after the attention layer (i.…
  8. Q9Suppose the input sentence is “tea you enjoy often”. Using the same matrices and processing method, what is the attenti…
  9. Q10Choose the correct representation of π for the given mask image.
  10. Q11How many permutations of π are possible for n = 5?
  11. Q12What percentage of entries in the attention matrix are zero (i.e., the sparsity)?
  12. Q13Consider the short sentence: “small models sometimes beat big ones” A model processes this sequence using the naive rel…
  13. Q14Consider the short sentence: “small models sometimes beat big ones” A model processes this sequence using the naive rel…
  14. Q15Consider the short sentence: “small models sometimes beat big ones” A model processes this sequence using the naive rel…
  15. Q16Consider the short sentence: “small models sometimes beat big ones” A model processes this sequence using the naive rel…
  16. Q17Consider the short sentence: “small models sometimes beat big ones” A model processes this sequence using the naive rel…