Question 1
A corpus containing the following five words is tokenized using WordPiece algorithm.
| Word | Frequency |
|---|---|
play | 3 |
played | 2 |
pray | 1 |
prey | 5 |
reply | 4 |
| Total | 15 |
Table 1: Word frequencies in the corpus before any merges
The initial character-level vocabulary is:
Based on the above data, answer the given subquestions.