← Tokens courseTOKENS · CHAPTER 6 OF 13 · 4 min READ · FULL COURSE

The context window: the model's desk

About access

Buyers receive the course-start link after checkout. This lesson link is unlisted, not account-based access control; anyone with the direct link can open it.

The story

Imagine a very bright assistant who can only work with papers on one desk. Anything on the desk, they can use. Anything not on the desk, they cannot see. The desk has a fixed size.

The one plain idea

The context window is the maximum number of tokens a model can look at in one go. It includes everything: your question, any documents you pasted, earlier turns of the chat, hidden instructions from the app, and the answer being written.

The kitchen-table picture

A kitchen table with a size limit. If you keep adding papers, something has to come off. Old papers slide to the floor. The assistant is no longer able to see them.

Facts and variation

  • Fact: every model has a context limit counted in tokens.
  • It varies: the size. Over the last few years, windows have grown from a few thousand tokens to hundreds of thousands, and some models advertise a million or more. Check the current number for the specific model you use.
  • Fact: input and output usually share the window. A very long document leaves less room for the answer.

The model has no memory between chats

By default, the model does not remember your earlier conversation. Apps give the feeling of memory by re-sending the earlier turns every time, which means they take up tokens on the desk. Some products add a separate memory feature that stores notes and inserts them into the window. Either way, what the model uses is whatever is on the desk right now.

Worked example

Illustrative: a 10,000-token window.

  • Hidden instructions from the app: 1,000
  • Pasted contract: 6,000
  • Earlier chat: 1,500
  • Your new question: 100
  • Left for the answer: 1,400 tokens, about 1,000 words.

If you paste a second contract of 4,000 tokens, the desk overflows. The app has to drop or shorten something.

Bigger is not automatically better

  • Cost. More tokens in, more tokens billed (Chapter 7).
  • Speed. Reading a very large input takes longer.
  • Attention is imperfect. Research and user experience have shown that models can be less reliable at using details buried in the middle of very long inputs than details near the start or end. How much it matters varies by model. A good habit is to put the most important instruction and facts where they are easy to see, and to give the model only what it needs.

Ways to deal with a small desk

  • Summarise older parts of a chat and keep the summary.
  • Retrieve: store documents elsewhere and bring only the relevant paragraphs onto the desk when needed. This idea is often called retrieval.
  • Split a big job into smaller steps.

No-code exercise

Take one long article you read this week. Estimate its words, then convert to tokens (words divided by 0.75). Would it fit in a 10,000-token window? Now estimate: if you wanted to ask it five questions, in a chat that re-sends everything each time, how many tokens would you have used in total?

Self-check

  1. What is a context window?
  2. What shares the window besides your question?
  3. Does the model remember last week's chat on its own?

Answers. 1) The maximum number of tokens the model can look at at once. 2) Documents, earlier turns, hidden instructions, and the answer. 3) Not by default. Memory is either re-sent text or a separate feature.

CURIOUS? TEST THE CLUES

Curiosity check

Pick an answer and see why. No scores, no pressure. All shop examples are invented practice scenarios.

01 What does a context limit constrain?
02 Is fitting a long document enough to guarantee perfect recall?
Course sources and freshness

Sources to read next

These are public, well-known references behind facts named in this course. Check them for current detail.

  • Vaswani et al., "Attention Is All You Need" (2017). The transformer paper.
  • Sennrich, Haddow, Birch, "Neural Machine Translation of Rare Words with Subword Units" (2015). Byte pair encoding for language models.
  • Philip Gage, "A New Algorithm for Data Compression" (1994). The original pair-merging idea.
  • Petrov et al., "Language Model Tokenizers Introduce Unfairness Between Languages" (NeurIPS 2023).
  • Radford et al., "Language Models are Unsupervised Multitask Learners" (GPT-2, 2019). Byte-level BPE and its 50,257-piece vocabulary.
  • Kudo and Richardson, "SentencePiece" (2018).
  • Rumbelow and Watkins, "SolidGoldMagikarp (plus, prompt generation)" (2023). Glitch tokens.
  • Your chosen provider's current tokenizer tool and price page, for live counts and prices.