Imagine you run a small shop. A customer writes: "Can I return this to the store?" The word return sounds clear, but you still need to know what this is. Is it an unopened mug bought yesterday? A made-to-order cake? A gift purchased at another branch? The sentence alone cannot tell you the right policy. Before you answer, you need more of the story.
Language works this way all the time. Consider these two sentences:
"The bank approved the shop's loan."
"The bank beside the market flooded."
The spelling bank is identical. In the first sentence, "loan" points toward a financial institution. In the second, "flooded" points toward the edge of a river or other water. If all you were given was "The bank...", you would not know which one was meant. That is not a failure of attention; the needed clue has not arrived.
A computer cannot simply keep one fixed meaning for bank and get both sentences right. It has to use the words around it. In a modern language model, text is divided into small pieces called tokens. A token may be a word, a piece of a word, or punctuation. The model handles patterns in these pieces and their surrounding context, rather than consulting a dictionary entry with one permanent meaning for every word.
Here is a commerce version. Your team's note says: "The charge was reversed." What happened? Perhaps a card payment was refunded. Perhaps a disputed payment became a chargeback. Perhaps someone is talking about reversing a delivery fee in an internal spreadsheet. You cannot send a confident customer reply yet. Ask which order and which kind of reversal, then check the payment record. The surrounding evidence matters more than the familiar sound of the sentence.
Notice the clue
Read: "The carton was light, but its shipping charge was heavy." The word light describes the carton as not weighing much; heavy describes a large cost, not a parcel you cannot lift. The topic and neighboring words steer the meaning. A person does this so naturally that it can feel invisible.
A Transformer uses a mechanism called self-attention to relate pieces of its input to one another. In the next lesson, we will make that idea concrete. For now, hold on to one simple point: the same word can need a different interpretation when the surrounding words change. This does not mean the model has human judgment or that its answers are automatically true. It can still miss a clue, misunderstand an ambiguous sentence, or state an invented detail with confidence.
Try it in your shop
A shopper says: "I need a replacement for the damaged one." What would you ask before promising anything? A good first question is, "Which item or order do you mean?" You might also need a photo and the store's actual replacement policy. Asking for the missing clue is more useful than guessing from a fluent sentence.
Try it yourself
Write two sentences with the same word used differently in a shop. Try charge, stock, or order. Circle the words that tell you which meaning is intended. If the clue is missing, add the question you would ask.
What to remember: Meaning depends on context. A good assistant, human or AI, should ask when the clue is missing.
Sources and accuracy notes
- Google Research's Transformer explanation uses bank and river as a word-sense example.
- Google's ML Crash Course explains tokens and self-attention.
- IBM's self-attention explainer describes how the mechanism relates pieces of an input.
The shop, customer, carton, and return examples are invented teaching examples, not real merchant policies.