The story
Imagine a very bright assistant who can only work with papers on one desk. Anything on the desk, they can use. Anything not on the desk, they cannot see. The desk has a fixed size.
The one plain idea
The context window is the maximum number of tokens a model can look at in one go. It includes everything: your question, any documents you pasted, earlier turns of the chat, hidden instructions from the app, and the answer being written.
The kitchen-table picture
A kitchen table with a size limit. If you keep adding papers, something has to come off. Old papers slide to the floor. The assistant is no longer able to see them.
Facts and variation
- Fact: every model has a context limit counted in tokens.
- It varies: the size. Over the last few years, windows have grown from a few thousand tokens to hundreds of thousands, and some models advertise a million or more. Check the current number for the specific model you use.
- Fact: input and output usually share the window. A very long document leaves less room for the answer.
The model has no memory between chats
By default, the model does not remember your earlier conversation. Apps give the feeling of memory by re-sending the earlier turns every time, which means they take up tokens on the desk. Some products add a separate memory feature that stores notes and inserts them into the window. Either way, what the model uses is whatever is on the desk right now.
Worked example
Illustrative: a 10,000-token window.
- Hidden instructions from the app: 1,000
- Pasted contract: 6,000
- Earlier chat: 1,500
- Your new question: 100
- Left for the answer: 1,400 tokens, about 1,000 words.
If you paste a second contract of 4,000 tokens, the desk overflows. The app has to drop or shorten something.
Bigger is not automatically better
- Cost. More tokens in, more tokens billed (Chapter 7).
- Speed. Reading a very large input takes longer.
- Attention is imperfect. Research and user experience have shown that models can be less reliable at using details buried in the middle of very long inputs than details near the start or end. How much it matters varies by model. A good habit is to put the most important instruction and facts where they are easy to see, and to give the model only what it needs.
Ways to deal with a small desk
- Summarise older parts of a chat and keep the summary.
- Retrieve: store documents elsewhere and bring only the relevant paragraphs onto the desk when needed. This idea is often called retrieval.
- Split a big job into smaller steps.
No-code exercise
Take one long article you read this week. Estimate its words, then convert to tokens (words divided by 0.75). Would it fit in a 10,000-token window? Now estimate: if you wanted to ask it five questions, in a chat that re-sends everything each time, how many tokens would you have used in total?
Self-check
- What is a context window?
- What shares the window besides your question?
- Does the model remember last week's chat on its own?
Answers. 1) The maximum number of tokens the model can look at at once. 2) Documents, earlier turns, hidden instructions, and the answer. 3) Not by default. Memory is either re-sent text or a separate feature.