The story
When you type into an AI chat and the answer appears word by word, that is not a special effect. That is literally how it works. The model produces one token, then the next, then the next.
The one plain idea
A language model is a next-token predictor. Given all the tokens so far, it works out how likely each possible next token is. Then one is chosen, added to the text, and the process repeats.
The kitchen-table picture
You know the game where someone starts a sentence and you finish it? "Peanut butter and ___." Most people say "jelly". You are predicting the next piece based on everything you have seen. A language model does this with an enormous list of possible next tokens, assigning each a probability.
Worked example
Text so far: "The capital of France is"
The model scores every token in its vocabulary. Illustrative numbers:
| Next token | Chance |
|---|---|
| " Paris" | 92% |
| " a" | 2% |
| " the" | 1% |
| everything else | 5% |
Usually " Paris" is picked. It is added to the text. Now the text is "The capital of France is Paris", and the model predicts again: probably "." Then perhaps it stops.
Picking from the list: the "temperature" dial
The model does not always pick the top token. Providers often let you control how adventurous the pick is.
- Low temperature: nearly always the most likely token. Steady and predictable. Good for facts and for code.
- High temperature: more willing to pick less likely tokens. More varied and creative, and also more likely to wander.
This is why asking the same question twice can give two different answers. Different dice rolls on the picks.
Why long answers cost more
Each new token needs the model to do its full computation again, looking back over what came before. So every token written has a real cost in computer time. This is the physical reason AI is billed by tokens (Chapter 7).
Stopping
The model has a special token that means "I am done," often called an end-of-sequence token. When it picks that, writing stops. You can also set a maximum number of tokens for an answer, which cuts it off if it runs long.
A common confusion
People say "the AI knows" or "the AI thinks." At the level of mechanics, it is repeatedly answering one question: given everything so far, what token probably comes next? The results can look like understanding, and they are useful. But knowing the mechanism helps you predict when it will be right and when it will confidently invent things. A smooth guess and a true fact look identical on the page.
No-code exercise
Ask an AI chat the same open question three times: "Give me a name for a coffee shop." Compare. Then ask it to "answer in exactly one word" three times. Where did the answers vary and where did they not? Write down what that tells you about picking from probabilities.
Self-check
- In one line, what does a language model do?
- What does a higher temperature do?
- Why does a longer answer cost more than a short one?
Answers. 1) Predicts the next token, over and over. 2) Makes it more willing to choose less likely tokens. 3) Each token needs its own round of computation.