Question 1
What is a primary motivation for using subword tokenization (like BPE or WordPiece) instead of word-level tokenization in transformers?
It increases vocabulary size exponentially
It completely eliminates the need for a tokenizer
It helps handle out-of-vocabulary words more effectively
It prevents the model from needing positional encoding