← Tokens courseTOKENS · CHAPTER 3 OF 13 · 4 min READ · FULL COURSE

How the chopping rules are learned

About access

Buyers receive the course-start link after checkout. This lesson link is unlisted, not account-based access control; anyone with the direct link can open it.

The story

Who decides that "ing" is one piece and "zqx" is not? Nobody sat down and wrote the list by hand. The list is learned from lots of text, by a simple trick that began as a way to shrink files.

In 1994, a programmer named Philip Gage described a data compression method that repeatedly replaces the most common pair of neighbouring symbols with one new symbol. Decades later, researchers (Sennrich, Haddow and Birch, 2015) adapted the same idea to build vocabularies for machine translation. It is called byte pair encoding, or BPE. Versions of it sit under many well-known language models.

The one plain idea

Start with the smallest pieces. Find the pair that appears next to each other most often. Glue it into one new piece. Repeat until the list is as long as you want.

The kitchen-table picture

Imagine a stack of recipe cards. You notice that "ch" and "ocolate" keep showing up side by side. You staple them together and write a new card: "chocolate". Then you notice "chocolate" and " chip" appear together a lot. Staple again. After many rounds, your most useful chunks are whole cards, and rare stuff stays in small bits.

Worked example (tiny and made up)

Training text: "low low low lower lowest"

  1. Start with letters: l o w, l o w, l o w, l o w e r, l o w e s t.
  2. The most common neighbouring pair is "l" + "o" (it appears 5 times). Glue it: "lo".
  3. Now "lo" + "w" is the most common pair. Glue it: "low".
  4. Now the pieces are: low, low, low, low + e + r, low + e + s + t.
  5. Next most common pair might be "e" + "r" or "e" + "s". Glue one, and so on.

After a few rounds, "low" is a single piece, and "er" and "est" appear as pieces too. The tokenizer learned that "low" is a useful chunk, and it learned that without anyone writing a rule.

What this means

Because the list is learned from data, it reflects the data. If the training text had lots of English web pages, English chunks become whole pieces. If it had little Swahili, Swahili words get chopped into many small bits. We return to this in Chapters 8 and 10.

A close cousin: bytes

Some tokenizers start from raw bytes, the basic computer units behind every character, instead of from letters. That way every possible text, including emoji and rare scripts, can be written down. Nothing is ever truly unknown. GPT-2 popularised this byte-level approach, and its vocabulary had 50,257 pieces.

Other well-known methods exist, such as WordPiece (used in BERT) and SentencePiece (a toolkit that works straight on raw text, including languages without spaces between words). You do not need to memorise them. The shared idea is the same: learn a good list of reusable pieces.

No-code exercise

Take this phrase: "banana bandana bandit". Do the glue game by hand. Write each word as letters. Find the most common neighbouring pair across all three. Glue it. Repeat three times. What pieces do you end up with?

Self-check

  1. Who wrote the list of tokens by hand?
  2. What is the one-line recipe for byte pair encoding?
  3. Why does the training text matter for the final tokenizer?

Answers. 1) Nobody. It is learned from text. 2) Repeatedly glue the most common neighbouring pair into a new piece. 3) The pieces that appear often become whole tokens, so the tokenizer favours what the text contained most.

CURIOUS? TEST THE CLUES

Curiosity check

Pick an answer and see why. No scores, no pressure. All shop examples are invented practice scenarios.

01 What does a BPE merge join?
02 Does the toy merge example reproduce every real tokenizer?
Course sources and freshness

Sources to read next

These are public, well-known references behind facts named in this course. Check them for current detail.

  • Vaswani et al., "Attention Is All You Need" (2017). The transformer paper.
  • Sennrich, Haddow, Birch, "Neural Machine Translation of Rare Words with Subword Units" (2015). Byte pair encoding for language models.
  • Philip Gage, "A New Algorithm for Data Compression" (1994). The original pair-merging idea.
  • Petrov et al., "Language Model Tokenizers Introduce Unfairness Between Languages" (NeurIPS 2023).
  • Radford et al., "Language Models are Unsupervised Multitask Learners" (GPT-2, 2019). Byte-level BPE and its 50,257-piece vocabulary.
  • Kudo and Richardson, "SentencePiece" (2018).
  • Rumbelow and Watkins, "SolidGoldMagikarp (plus, prompt generation)" (2023). Glitch tokens.
  • Your chosen provider's current tokenizer tool and price page, for live counts and prices.