Opening the paper…
Large Language Models, Quiz 2
What is a primary motivation for using subword tokenization (like BPE or WordPiece) instead of word-level tokenization in transformers?
What is a primary motivation for using subword tokenization (like BPE or WordPiece) instead of word-level tokenization in transformers? What is a key implication of scaling laws in large language models? Consider the Following Assertion and Reason pair :\ **Assertion :** In Byte Pair Encoding (BPE), each merge operation adds exactly one new token to the vocabulary.\ **Reason :** BPE helps reduce the size of the initial vocabulary by merging the most frequent tokens.