The story
An ID number like 1,234 tells the model nothing. It is just a label, like a locker number. Locker 1,234 is not "close" to locker 1,235 in meaning. So how does a model get from a label to meaning?
The one plain idea
Each token ID is swapped for a long list of numbers, called an embedding. The model learns those numbers during training, so tokens with similar uses end up with similar lists.
The kitchen-table picture
Think of a map. Every town has coordinates. Towns near each other have similar coordinates. Now imagine a map with not two coordinates but hundreds or thousands. Every token is a town on this huge map. "Coffee" and "tea" sit close together. "Coffee" and "invoice" sit far apart. Nobody drew the map. The model drew it while learning from text, placing tokens near others that appear in similar situations.
The list of numbers for one token is its address on the meaning map. The length of the list varies by model.
Worked example (simplified to 2 numbers)
Illustrative only. Real embeddings are far longer.
| Token | Number 1 | Number 2 |
|---|---|---|
| coffee | 0.9 | 0.8 |
| tea | 0.8 | 0.9 |
| invoice | -0.7 | 0.1 |
"coffee" and "tea" are close. "invoice" is far from both. A model can now do something useful: when the sentence says "I'd like a hot ___", it leans towards tokens near "coffee" and "tea", not "invoice".
Position matters too
Words in a different order mean different things: "dog bites man" is not "man bites dog". So the model also gets information about where each token sits in the sequence. Different models do this in different ways. The result is the same: each token arrives carrying "what I am" and "where I am".
The transformer in one paragraph
Most well-known language models today are built on an idea called the transformer, introduced in the 2017 paper "Attention Is All You Need." Its key move is attention: for each token, the model looks at the other tokens in the text and decides which ones matter most for understanding this one. In "The bag didn't fit in the car because it was too big," attention helps work out what "it" points to. The input to all of this is the stream of token embeddings from this chapter.
What this means for you
Tokens are the doorway. Text goes in as tokens, becomes lists of numbers, and the model's maths happens on those numbers. The quality of this doorway, meaning how well the pieces are chosen, affects everything after it.
No-code exercise
Draw a 10 by 10 grid on paper. Place these words where you think they belong so that similar words are close: coffee, tea, price, discount, sale, cat, dog, warehouse, truck, delivery. Then check: did you put "truck" next to "delivery"? That is what embeddings capture.
Self-check
- Why is a token ID alone not enough for the model?
- What is an embedding?
- Why does the model also need to know position?
Answers. 1) It is just a label with no meaning built in. 2) A learned list of numbers that acts as the token's address on a meaning map. 3) Order changes meaning.