Quiz Space

Deep Learning Practice Quiz 1: 23 February 2025 (January 2025 term)

Question 1

+3 marksOne correct option
  1. A

    ["hugging", "face", "is", "awesome"]

  2. B

    [102, 463, 509, 101, 2020]

  3. C

    ["hugging", "face", "is", "awesome", ""]

  4. D

    "Hugging Face is awesome!"

Question 2

+3 marksOne correct option

Consider the following code where a tokenizer is created using the BPE model:

python
from tokenizers import Tokenizer
from tokenizers.models import BPE
from tokenizers.trainers import BpeTrainer
from tokenizers.normalizers import Lowercase
from tokenizers.pre_tokenizers import Whitespace
data = ["existent", "non", "exist", "word"]
# Create a tokenizer with BPE model without specifying the unk_token
model = BPE()
tokenizer = Tokenizer(model)
# Normalizer and Pre-tokenizer
tokenizer.normalizer = Lowercase()
tokenizer.pre_tokenizer = Whitespace()
# Train the tokenizer
trainer = BpeTrainer(vocab_size=5000, special_tokens=["<s>", "</s>", "<pad>"])
tokenizer.train_from_iterator(data, trainer)
# Access the tokenizer's vocabulary
vocab = tokenizer.get_vocab()
# Test tokenization with a word not present in the vocabulary
tokens = tokenizer.encode("nonexistentword").tokens
print("Tokens:", tokens)

Which of the following statements is correct about the tokenization output when the word "nonexistentword" is encountered?

  1. A

    The tokenizer will output the token [UNK] because "nonexistentword" is not in the vocabulary.

  2. B

    The tokenizer will output the word "nonexistentword" as a single token.

  3. C

    The tokenizer will output "non", "existent", "word" as separate tokens since it has split the word into known subwords.

  4. D

    The tokenizer will output "non", "exist", "ent", "word" as separate tokens since it has split the word into known subwords.

  5. E

    The tokenizer will output an error because "nonexistentword" is not present in the vocabulary and the unk_token was not defined.

Question 3

+3 marksOne correct option

Consider the following code snippet:

python
from tokenizers import Tokenizer
from tokenizers.models import BPE
from tokenizers.trainers import BpeTrainer
from tokenizers.normalizers import Lowercase
from tokenizers.pre_tokenizers import Whitespace
from tokenizers.processors import TemplateProcessing
data = ["hello", "world"]
# Create a tokenizer with BPE model
model = BPE()
tokenizer = Tokenizer(model)
# Normalizer and Pre-tokenizer
tokenizer.normalizer = Lowercase()
tokenizer.pre_tokenizer = Whitespace()
# Trainer for the tokenizer
trainer = BpeTrainer(vocab_size=5000, special_tokens=["<s>", "</s>", "<pad>", "<unk>"])
tokenizer.train_from_iterator(data, trainer)
# Post-processor
tokenizer.post_processor = TemplateProcessing(single="[CLS] $0 [SEP]",
special_tokens=[("[CLS]", 2), ("[SEP]", 3)])
# Encode input and get the token IDs
encoded = tokenizer.encode("Hello world")
print("Token IDs:", encoded.ids)

What will likely be printed as the output?

  1. A

    [2, 3]

  2. B

    [3, 0, 2]

  3. C

    [2, 0, 1, 3]

  4. D

    [3, 1, 0, 2]

13 more questions in this paper

Sign in with Google — it is free — to see every question with its answer and explanation, practise it in learning mode, or take it as a timed mock test.

More on the Deep Learning Practice Quiz 1 23 Feb 2025 paper

The IIT Madras BS Deep Learning Practice (Deep Learning Practice) Quiz 1 paper sat on 23 Feb 2025, in the January 2025 term: 16 questions for 50 marks in 120 minutes. The first 3 questions are below. Sign in with Google — it is free — to see the whole paper with its answers and explanations, in learning mode or as a timed mock test.

FeatureDeep Learning Practice Quiz 1 23 Feb 2025 at a glance
TermJanuary 2025 term
SubjectDeep Learning Practice
Course codeBSDA5013
Questions16
Marks50
Duration120 min
MCQ9
MSQ4
Numerical3
Official paperIIT M DEGREE AN EXAM QDB2 23 Feb 2025
Negative markingNo negative marking.
Updated

Same Quiz 1, other subjects

More Deep Learning Practice