← Tokens courseTOKENS · CHAPTER 7 OF 13 · 4 min READ · FREE

Tokens as money: how AI is priced

The story

A shop sells electricity by the kilowatt-hour. A taxi sells rides by the mile. AI providers mostly sell language models by the token. It is the meter on the wall.

The one plain idea

When you use a model through a paid service or an API, you are usually billed for the tokens you send in and the tokens you get out, often at different rates.

The kitchen-table picture

A print shop charges per page. But it charges one rate for pages you hand in to be scanned, and a higher rate for pages it prints for you. Same idea: input tokens and output tokens.

Facts and variation

  • Fact: pricing is commonly quoted per million tokens.
  • Common pattern (it varies): output tokens cost more than input tokens, since writing is the expensive part (Chapter 5).
  • It varies: bigger, more capable models cost more per token than smaller ones. Some providers discount repeated input that they have seen before, often called caching. Some models spend extra tokens "thinking" before answering, and those hidden tokens may be billed too. Check the provider's price page.
  • Consumer chat apps are often a flat monthly fee, with usage limits behind the scenes that are also measured in tokens or something like them.

Worked example

All prices below are made-up round numbers for teaching. They are not any real provider's prices.

Say a model costs $2 per million input tokens and $8 per million output tokens.

A customer-service bot answers one question:

  • Input: 1,500 tokens (instructions + the customer's message + order history)
  • Output: 300 tokens

Cost = (1,500 / 1,000,000) x $2 + (300 / 1,000,000) x $8 = $0.003 + $0.0024 = $0.0054, about half a cent.

Half a cent sounds like nothing. Now scale: 200,000 conversations a month. 200,000 x $0.0054 = $1,080 a month.

Now suppose the team lets the bot re-send the whole chat history each turn, and conversations average 6 turns, so input grows. Input might triple. The monthly bill is no longer $1,080. Run the numbers again: input cost goes from $0.003 to $0.009 per question, so the total per question is about $0.0114, and the month is about $2,280. Small design choices change the bill a lot.

Tokens per second: the speed meter

You will also hear "tokens per second." That is how fast a model writes. A faster model feels snappier, and for a business the tokens per second per computer chip decides how many customers one machine can serve. This is where the technical meter and the money meter meet.

Commerce angle

If you run an online store and put an AI helper on it, three numbers drive your cost: tokens in per question, tokens out per question, and number of questions. Shortening the instructions, trimming product data you send, and capping answer length are the main levers. Pick the cheapest model that does the job well enough, and test that on real examples.

No-code exercise

Build a small sheet with three columns: input tokens, output tokens, number of requests per month. Use the example prices above. Make three scenarios: a small shop (1,000 requests), a mid-size shop (50,000), a big one (2,000,000). Which scenario would make you think hard about model choice?

Self-check

  1. What two kinds of tokens are billed?
  2. Why are output tokens often priced higher?
  3. Name two ways to lower a token bill.

Answers. 1) Input and output. 2) Producing each one takes a separate round of computation. 3) Shorter instructions or less data in, capped answers, a smaller model, caching where offered.

What comes next (paid)

You now know what a token costs. The rest of the course shows why the same sentence can cost about 3 times more in some languages, why models fail at counting letters, and how to cut your own bill.

CURIOUS? TEST THE CLUES

Curiosity check

Pick an answer and see why. No scores, no pressure. All shop examples are invented practice scenarios.

01 At an illustrative $2 per million input tokens, what do 5,000 input tokens cost?
02 What must you check before estimating a real bill?
Course sources and freshness

Sources to read next

These are public, well-known references behind facts named in this course. Check them for current detail.

  • Vaswani et al., "Attention Is All You Need" (2017). The transformer paper.
  • Sennrich, Haddow, Birch, "Neural Machine Translation of Rare Words with Subword Units" (2015). Byte pair encoding for language models.
  • Philip Gage, "A New Algorithm for Data Compression" (1994). The original pair-merging idea.
  • Petrov et al., "Language Model Tokenizers Introduce Unfairness Between Languages" (NeurIPS 2023).
  • Radford et al., "Language Models are Unsupervised Multitask Learners" (GPT-2, 2019). Byte-level BPE and its 50,257-piece vocabulary.
  • Kudo and Richardson, "SentencePiece" (2018).
  • Rumbelow and Watkins, "SolidGoldMagikarp (plus, prompt generation)" (2023). Glitch tokens.
  • Your chosen provider's current tokenizer tool and price page, for live counts and prices.