← Tokens courseTOKENS · CHAPTER 5 OF 13 · 4 min READ · FULL COURSE

How a model writes: one token at a time

About access

Buyers receive the course-start link after checkout. This lesson link is unlisted, not account-based access control; anyone with the direct link can open it.

The story

When you type into an AI chat and the answer appears word by word, that is not a special effect. That is literally how it works. The model produces one token, then the next, then the next.

The one plain idea

A language model is a next-token predictor. Given all the tokens so far, it works out how likely each possible next token is. Then one is chosen, added to the text, and the process repeats.

The kitchen-table picture

You know the game where someone starts a sentence and you finish it? "Peanut butter and ___." Most people say "jelly". You are predicting the next piece based on everything you have seen. A language model does this with an enormous list of possible next tokens, assigning each a probability.

Worked example

Text so far: "The capital of France is"

The model scores every token in its vocabulary. Illustrative numbers:

Next tokenChance
" Paris"92%
" a"2%
" the"1%
everything else5%

Usually " Paris" is picked. It is added to the text. Now the text is "The capital of France is Paris", and the model predicts again: probably "." Then perhaps it stops.

Picking from the list: the "temperature" dial

The model does not always pick the top token. Providers often let you control how adventurous the pick is.

  • Low temperature: nearly always the most likely token. Steady and predictable. Good for facts and for code.
  • High temperature: more willing to pick less likely tokens. More varied and creative, and also more likely to wander.

This is why asking the same question twice can give two different answers. Different dice rolls on the picks.

Why long answers cost more

Each new token needs the model to do its full computation again, looking back over what came before. So every token written has a real cost in computer time. This is the physical reason AI is billed by tokens (Chapter 7).

Stopping

The model has a special token that means "I am done," often called an end-of-sequence token. When it picks that, writing stops. You can also set a maximum number of tokens for an answer, which cuts it off if it runs long.

A common confusion

People say "the AI knows" or "the AI thinks." At the level of mechanics, it is repeatedly answering one question: given everything so far, what token probably comes next? The results can look like understanding, and they are useful. But knowing the mechanism helps you predict when it will be right and when it will confidently invent things. A smooth guess and a true fact look identical on the page.

No-code exercise

Ask an AI chat the same open question three times: "Give me a name for a coffee shop." Compare. Then ask it to "answer in exactly one word" three times. Where did the answers vary and where did they not? Write down what that tells you about picking from probabilities.

Self-check

  1. In one line, what does a language model do?
  2. What does a higher temperature do?
  3. Why does a longer answer cost more than a short one?

Answers. 1) Predicts the next token, over and over. 2) Makes it more willing to choose less likely tokens. 3) Each token needs its own round of computation.

CURIOUS? TEST THE CLUES

Curiosity check

Pick an answer and see why. No scores, no pressure. All shop examples are invented practice scenarios.

01 What does next-token prediction choose?
02 Does higher temperature guarantee better answers?
Course sources and freshness

Sources to read next

These are public, well-known references behind facts named in this course. Check them for current detail.

  • Vaswani et al., "Attention Is All You Need" (2017). The transformer paper.
  • Sennrich, Haddow, Birch, "Neural Machine Translation of Rare Words with Subword Units" (2015). Byte pair encoding for language models.
  • Philip Gage, "A New Algorithm for Data Compression" (1994). The original pair-merging idea.
  • Petrov et al., "Language Model Tokenizers Introduce Unfairness Between Languages" (NeurIPS 2023).
  • Radford et al., "Language Models are Unsupervised Multitask Learners" (GPT-2, 2019). Byte-level BPE and its 50,257-piece vocabulary.
  • Kudo and Richardson, "SentencePiece" (2018).
  • Rumbelow and Watkins, "SolidGoldMagikarp (plus, prompt generation)" (2023). Glitch tokens.
  • Your chosen provider's current tokenizer tool and price page, for live counts and prices.