The story
Who decides that "ing" is one piece and "zqx" is not? Nobody sat down and wrote the list by hand. The list is learned from lots of text, by a simple trick that began as a way to shrink files.
In 1994, a programmer named Philip Gage described a data compression method that repeatedly replaces the most common pair of neighbouring symbols with one new symbol. Decades later, researchers (Sennrich, Haddow and Birch, 2015) adapted the same idea to build vocabularies for machine translation. It is called byte pair encoding, or BPE. Versions of it sit under many well-known language models.
The one plain idea
Start with the smallest pieces. Find the pair that appears next to each other most often. Glue it into one new piece. Repeat until the list is as long as you want.
The kitchen-table picture
Imagine a stack of recipe cards. You notice that "ch" and "ocolate" keep showing up side by side. You staple them together and write a new card: "chocolate". Then you notice "chocolate" and " chip" appear together a lot. Staple again. After many rounds, your most useful chunks are whole cards, and rare stuff stays in small bits.
Worked example (tiny and made up)
Training text: "low low low lower lowest"
- Start with letters: l o w, l o w, l o w, l o w e r, l o w e s t.
- The most common neighbouring pair is "l" + "o" (it appears 5 times). Glue it: "lo".
- Now "lo" + "w" is the most common pair. Glue it: "low".
- Now the pieces are: low, low, low, low + e + r, low + e + s + t.
- Next most common pair might be "e" + "r" or "e" + "s". Glue one, and so on.
After a few rounds, "low" is a single piece, and "er" and "est" appear as pieces too. The tokenizer learned that "low" is a useful chunk, and it learned that without anyone writing a rule.
What this means
Because the list is learned from data, it reflects the data. If the training text had lots of English web pages, English chunks become whole pieces. If it had little Swahili, Swahili words get chopped into many small bits. We return to this in Chapters 8 and 10.
A close cousin: bytes
Some tokenizers start from raw bytes, the basic computer units behind every character, instead of from letters. That way every possible text, including emoji and rare scripts, can be written down. Nothing is ever truly unknown. GPT-2 popularised this byte-level approach, and its vocabulary had 50,257 pieces.
Other well-known methods exist, such as WordPiece (used in BERT) and SentencePiece (a toolkit that works straight on raw text, including languages without spaces between words). You do not need to memorise them. The shared idea is the same: learn a good list of reusable pieces.
No-code exercise
Take this phrase: "banana bandana bandit". Do the glue game by hand. Write each word as letters. Find the most common neighbouring pair across all three. Glue it. Repeat three times. What pieces do you end up with?
Self-check
- Who wrote the list of tokens by hand?
- What is the one-line recipe for byte pair encoding?
- Why does the training text matter for the final tokenizer?
Answers. 1) Nobody. It is learned from text. 2) Repeatedly glue the most common neighbouring pair into a new piece. 3) The pieces that appear often become whole tokens, so the tokenizer favours what the text contained most.