Large Language Models, Quiz 2
A corpus containing the following five words is tokenized using WordPiece algorithm.
| Word | Frequency |
|---|---|
play | 3 |
played | 2 |
pray | 1 |
prey | 5 |
reply | 4 |
| Total | 15 |
Table 1: Word frequencies in the corpus before any merges
The initial character-level vocabulary is:
Based on the above data, answer the given subquestions.
A corpus containing the following five words is tokenized using WordPiece algorithm. | Word | Frequency | |---|---| | `play` $< w >$ | 3 | | `played` $< w >$ | 2 | | `pray` $< w >$ | 1 | | `prey` $< w >$ | 5 | | `reply` $< w >$ | 4 | | **Total** | **15** | Table 1: Word frequencies in the corpus before any merges The initial character-level vocabulary is: $$V_0 = \{\text{‘p’}, \text{‘r’}, \text{‘l’}, \text{‘a’}, \text{‘y’}, \text{‘e’}, \text{‘d’}, \text{‘i’}, \text{‘n’}, \text{‘g’}, < w >\}.$$ Based on the above data, answer the given subquestions. Figure from the original question paper A corpus containing the following five words is tokenized using WordPiece algorithm. | Word | Frequency | |---|---| | `play` $< w >$ | 3 | | `played` $< w >$ | 2 | | `pray` $< w >$ | 1 | | `prey` $< w >$ | 5 | | `reply` $< w >$ | 4 | | **Total** | **15** | Table 1: Word frequencies in the corpus before any merges The initial character-level vocabulary is: $$V_0 = \{\text{‘p’}, \text{‘r’}, \text{‘l’}, \text{‘a’}, \text{‘y’}, \text{‘e’}, \text{‘d’}, \text{‘i’}, \text{‘n’}, \text{‘g’}, < w >\}.$$ Based on the above data, answer the given subquestions. Figure from the original question paper A corpus containing the following five words is tokenized using WordPiece algorithm. | Word | Frequency | |---|---| | `play` $< w >$ | 3 | | `played` $< w >$ | 2 | | `pray` $< w >$ | 1 | | `prey` $< w >$ | 5 | | `reply` $< w >$ | 4 | | **Total** | **15** | Table 1: Word frequencies in the corpus before any merges The initial character-level vocabulary is: $$V_0 = \{\text{‘p’}, \text{‘r’}, \text{‘l’}, \text{‘a’}, \text{‘y’}, \text{‘e’}, \text{‘d’}, \text{‘i’}, \text{‘n’}, \text{‘g’}, < w >\}.$$ Based on the above data, answer the given subquestions. Which token will be generated in the first merge?