Quiz Space

Large Language Models · Quiz 2 · 1 Dec 2024 · September 2024 term

Question 13: Which of the following components in the data pre-proces…

Question 13

+3 marksOne correct option

Which of the following components in the data pre-processing pipeline removes pages that contain bad words?

  1. A

    Language Identification

  2. B

    Exact Deduplication

  3. C

    Fuzzy Deduplication

  4. D

    ML classifiers for quality filtering

  5. E

    Simple heuristics to detect toxic contents

Show answer

Correct answer

  • E

    Simple heuristics to detect toxic contents

Question 13 of 16 in the IIT Madras BS Large Language Models (LLM) Quiz 2 paper sat on 1 Dec 2024, in the September 2024 term (IIT M DEGREE AN EXAM QDB2 01 Dec 2024). It carries 3 marks.

This question was also asked in

More questions from this paper

  1. Q1Consider the following dictionary with the number of word occurrences in a corpus: Note: Append identifier/special symb…
  2. Q2Consider the following dictionary with the number of word occurrences in a corpus: Note: Append identifier/special symb…
  3. Q3Consider the following dictionary with the number of word occurrences in a corpus: Note: Append identifier/special symb…
  4. Q4Consider the following dictionary with the number of word occurrences in a corpus: Note: Append identifier/special symb…
  5. Q5Consider the following dictionary with the number of word occurrences in a corpus: Note: Append identifier/special symb…
  6. Q6Consider the following dictionary with the number of word occurrences in a corpus: Note: Append identifier/special symb…
  7. Q7Consider the following dictionary with the number of word occurrences in a corpus: Note: Identifier/special symbol <…
  8. Q8Consider the following dictionary with the number of word occurrences in a corpus: Note: Identifier/special symbol <…
  9. Q9Which of the following represent(s) the normalization step(s) in building a tokenizer for the English language?
  10. Q10Choose all the aspects of the pre-training datasets that impact the model’s performance.
  11. Q11The strikeout words in the passage given below denote the words to be dropped from the original sentence. “Metacognitio…
  12. Q12Suppose you are working on prefix language modeling. The sequence length is 32 and the first two tokens represent the t…
  13. Q14Which of the following is an important implication of the scaling law:
  14. Q15Which of the following are the design choices for building a large language model?
  15. Q16Consider an ideal data preprocessing pipeline to prepare a dataset. Which of the following sentences or sets of paragra…