Deep Learning Practice, Quiz 1
Consider subword tokenization methods such as Byte Pair Encoding (BPE) and WordPiece as employed in transformer-based language modeling pipelines. Which of the following statements correctly describes their fundamental properties?
Consider subword tokenization methods such as Byte Pair Encoding (BPE) and WordPiece as employed in transformer-based language modeling pipelines. Which of the following statements correctly describes their fundamental properties? Consider two subword tokenizers trained on the same text corpus. Tokenizer A uses a vocabulary of 8,000 tokens, whereas Tokenizer B uses a vocabulary of 64,000 tokens. Which of the following statements most accurately characterizes the implications of these vocabulary sizes for training efficiency in transformer-based language models? A subword tokenizer decomposes rare chemical entity names into a large number of fragments. During downstream fine-tuning, the model exhibits degraded performance on a chemical named entity recognition (NER) task. Which of the following represents the most principled corrective action?