← Tokens courseTOKENS · CHAPTER 4 OF 13 · 4 min READ · FULL COURSE

From tokens to numbers to meaning

About access

Buyers receive the course-start link after checkout. This lesson link is unlisted, not account-based access control; anyone with the direct link can open it.

The story

An ID number like 1,234 tells the model nothing. It is just a label, like a locker number. Locker 1,234 is not "close" to locker 1,235 in meaning. So how does a model get from a label to meaning?

The one plain idea

Each token ID is swapped for a long list of numbers, called an embedding. The model learns those numbers during training, so tokens with similar uses end up with similar lists.

The kitchen-table picture

Think of a map. Every town has coordinates. Towns near each other have similar coordinates. Now imagine a map with not two coordinates but hundreds or thousands. Every token is a town on this huge map. "Coffee" and "tea" sit close together. "Coffee" and "invoice" sit far apart. Nobody drew the map. The model drew it while learning from text, placing tokens near others that appear in similar situations.

The list of numbers for one token is its address on the meaning map. The length of the list varies by model.

Worked example (simplified to 2 numbers)

Illustrative only. Real embeddings are far longer.

TokenNumber 1Number 2
coffee0.90.8
tea0.80.9
invoice-0.70.1

"coffee" and "tea" are close. "invoice" is far from both. A model can now do something useful: when the sentence says "I'd like a hot ___", it leans towards tokens near "coffee" and "tea", not "invoice".

Position matters too

Words in a different order mean different things: "dog bites man" is not "man bites dog". So the model also gets information about where each token sits in the sequence. Different models do this in different ways. The result is the same: each token arrives carrying "what I am" and "where I am".

The transformer in one paragraph

Most well-known language models today are built on an idea called the transformer, introduced in the 2017 paper "Attention Is All You Need." Its key move is attention: for each token, the model looks at the other tokens in the text and decides which ones matter most for understanding this one. In "The bag didn't fit in the car because it was too big," attention helps work out what "it" points to. The input to all of this is the stream of token embeddings from this chapter.

What this means for you

Tokens are the doorway. Text goes in as tokens, becomes lists of numbers, and the model's maths happens on those numbers. The quality of this doorway, meaning how well the pieces are chosen, affects everything after it.

No-code exercise

Draw a 10 by 10 grid on paper. Place these words where you think they belong so that similar words are close: coffee, tea, price, discount, sale, cat, dog, warehouse, truck, delivery. Then check: did you put "truck" next to "delivery"? That is what embeddings capture.

Self-check

  1. Why is a token ID alone not enough for the model?
  2. What is an embedding?
  3. Why does the model also need to know position?

Answers. 1) It is just a label with no meaning built in. 2) A learned list of numbers that acts as the token's address on a meaning map. 3) Order changes meaning.

CURIOUS? TEST THE CLUES

Curiosity check

Pick an answer and see why. No scores, no pressure. All shop examples are invented practice scenarios.

01 What is an embedding?
02 Do similar vectors prove a claim is true?
Course sources and freshness

Sources to read next

These are public, well-known references behind facts named in this course. Check them for current detail.

  • Vaswani et al., "Attention Is All You Need" (2017). The transformer paper.
  • Sennrich, Haddow, Birch, "Neural Machine Translation of Rare Words with Subword Units" (2015). Byte pair encoding for language models.
  • Philip Gage, "A New Algorithm for Data Compression" (1994). The original pair-merging idea.
  • Petrov et al., "Language Model Tokenizers Introduce Unfairness Between Languages" (NeurIPS 2023).
  • Radford et al., "Language Models are Unsupervised Multitask Learners" (GPT-2, 2019). Byte-level BPE and its 50,257-piece vocabulary.
  • Kudo and Richardson, "SentencePiece" (2018).
  • Rumbelow and Watkins, "SolidGoldMagikarp (plus, prompt generation)" (2023). Glitch tokens.
  • Your chosen provider's current tokenizer tool and price page, for live counts and prices.