The story
Imagine you run a warehouse that ships boxes. A box can hold only so much. If a customer sends you a huge sofa, you cannot ship it whole. You take it apart into pieces that fit boxes. At the other end, someone puts it back together.
A tokenizer does the same with text. It is a small program that sits in front of the model. Its job: take your text, cut it into tokens, and hand each token a number. When the model answers, the tokenizer turns numbers back into text.
The one plain idea
The tokenizer is a fixed lookup book. It has a list of allowed pieces, called the vocabulary. Each piece has an ID number. The tokenizer cuts your text into pieces from that list. The model never sees your letters. It sees only ID numbers.
The kitchen-table picture
Think of a phone book, but for text pieces instead of people. "the" is entry 1,234 (made-up number). "ing" is entry 5,678. The model only ever reads the numbers. It has never seen the word "the" in its life. It has only seen 1,234.
Why not just use whole words?
You could build a vocabulary of every word. It sounds simple. It fails for three reasons:
- Too many words. Names, typos, brand names, new slang, and every language together make millions of forms. The book would be enormous.
- New words appear daily. A brand name invented tomorrow would not be in the book. The model would have no way to read it.
- Related words look unrelated. "run", "running", "runner" would be three unrelated numbers, with nothing connecting them.
Why not just use single letters?
You could also use only letters. Then nothing is ever unknown. But text becomes very long. A sentence of 50 characters needs 50 pieces. Models handle long sequences slowly and expensively, and they have a limit on how much they can hold at once (Chapter 6).
The middle path: sub-words
Modern models use sub-words: pieces between letters and words. Common words get their own piece. Rare words are built from smaller pieces. Nothing is ever fully unknown, and text stays reasonably short. This is the deal almost every modern language model makes.
Worked example
The made-up word "unshoppable". A whole-word vocabulary would fail: not in the book. A sub-word tokenizer might split it into "un", "shop", "pable". The model has seen those pieces in other words, so it can still make a good guess at the meaning. This is called handling the out-of-vocabulary problem, and sub-words are the standard fix.
Different models, different books
Each model family usually comes with its own tokenizer and its own vocabulary. Vocabularies commonly run from tens of thousands to a few hundred thousand pieces, and it varies. This matters in one practical way: the same sentence can have a different token count on different models. Do not assume one model's count carries over to another.
No-code exercise
Pick one sentence with a made-up word, like "The shopbot was unscrollable." Run it through a public tokenizer demo. Note which pieces it chose. Does it split where you would?
Self-check
- What does a tokenizer do?
- What is a vocabulary?
- Why not use only whole words? Name two reasons.
- Will two different models always count the same tokens for the same sentence?
Answers. 1) Cuts text into tokens and maps each to an ID number, and back again. 2) The fixed list of pieces the tokenizer is allowed to use. 3) Too many words; new words keep appearing; related words look unrelated. 4) No. It varies by tokenizer.