Question 1
["hugging", "face", "is", "awesome"]
[102, 463, 509, 101, 2020]
["hugging", "face", "is", "awesome", ""]
"Hugging Face is awesome!"
["hugging", "face", "is", "awesome"]
[102, 463, 509, 101, 2020]
["hugging", "face", "is", "awesome", ""]
"Hugging Face is awesome!"
Consider the following code where a tokenizer is created using the BPE model:
from tokenizers import Tokenizerfrom tokenizers.models import BPEfrom tokenizers.trainers import BpeTrainerfrom tokenizers.normalizers import Lowercasefrom tokenizers.pre_tokenizers import Whitespace
data = ["existent", "non", "exist", "word"]
# Create a tokenizer with BPE model without specifying the unk_tokenmodel = BPE()tokenizer = Tokenizer(model)
# Normalizer and Pre-tokenizertokenizer.normalizer = Lowercase()tokenizer.pre_tokenizer = Whitespace()
# Train the tokenizertrainer = BpeTrainer(vocab_size=5000, special_tokens=["<s>", "</s>", "<pad>"])tokenizer.train_from_iterator(data, trainer)
# Access the tokenizer's vocabularyvocab = tokenizer.get_vocab()
# Test tokenization with a word not present in the vocabulary
tokens = tokenizer.encode("nonexistentword").tokensprint("Tokens:", tokens)Which of the following statements is correct about the tokenization output when the word "nonexistentword" is encountered?
The tokenizer will output the token [UNK] because "nonexistentword" is not in the vocabulary.
The tokenizer will output the word "nonexistentword" as a single token.
The tokenizer will output "non", "existent", "word" as separate tokens since it has split the word into known subwords.
The tokenizer will output "non", "exist", "ent", "word" as separate tokens since it has split the word into known subwords.
The tokenizer will output an error because "nonexistentword" is not present in the vocabulary and the unk_token was not defined.
Consider the following code snippet:
from tokenizers import Tokenizerfrom tokenizers.models import BPEfrom tokenizers.trainers import BpeTrainerfrom tokenizers.normalizers import Lowercasefrom tokenizers.pre_tokenizers import Whitespacefrom tokenizers.processors import TemplateProcessing
data = ["hello", "world"]
# Create a tokenizer with BPE modelmodel = BPE()tokenizer = Tokenizer(model)
# Normalizer and Pre-tokenizertokenizer.normalizer = Lowercase()tokenizer.pre_tokenizer = Whitespace()
# Trainer for the tokenizertrainer = BpeTrainer(vocab_size=5000, special_tokens=["<s>", "</s>", "<pad>", "<unk>"])tokenizer.train_from_iterator(data, trainer)
# Post-processortokenizer.post_processor = TemplateProcessing(single="[CLS] $0 [SEP]", special_tokens=[("[CLS]", 2), ("[SEP]", 3)])
# Encode input and get the token IDsencoded = tokenizer.encode("Hello world")print("Token IDs:", encoded.ids)What will likely be printed as the output?
[2, 3]
[3, 0, 2]
[2, 0, 1, 3]
[3, 1, 0, 2]
Sign in with Google — it is free — to see every question with its answer and explanation, practise it in learning mode, or take it as a timed mock test.
The IIT Madras BS Deep Learning Practice (Deep Learning Practice) Quiz 1 paper sat on 23 Feb 2025, in the January 2025 term: 16 questions for 50 marks in 120 minutes. The first 3 questions are below. Sign in with Google — it is free — to see the whole paper with its answers and explanations, in learning mode or as a timed mock test.
| Feature | Deep Learning Practice Quiz 1 23 Feb 2025 at a glance |
|---|---|
| Term | January 2025 term |
| Subject | Deep Learning Practice |
| Course code | BSDA5013 |
| Questions | 16 |
| Marks | 50 |
| Duration | 120 min |
| MCQ | 9 |
| MSQ | 4 |
| Numerical | 3 |
| Official paper | IIT M DEGREE AN EXAM QDB2 23 Feb 2025 |
| Negative marking | No negative marking. |
| Updated |