A word becomes a place
A model cannot read letters. Every word it takes in is turned into a list of numbers — a point in a huge space. The trick that makes this useful: words that behave alike end up close together in that space.
Click a word on the map. The lines show its nearest neighbors — the words the model has learned to treat as its relatives.
Every word carries a running notebook
As the sentence travels through the model, each word keeps one running list of numbers — its own notebook of what it means so far. Every layer reads the notebooks and adds notes. Nothing is ever erased; the model only ever adds.
Click each stage of the trip. (The real name for the notebook is the residual stream — one of only five terms this page will ask you to keep.)
Attention: each word gathers its context
The first pass inside the model is attention. Each word looks back at the words before it and copies a little from each — the lines below, with thickness meaning how much gets copied.
Click the word “it”. Who does it belong to? The model settles that by copying hardest from one earlier word.
Why a word may never look ahead
The toggle in Block 3 was not decoration. A language model is trained for one job: given everything so far, predict the next word. That job is only honest if the word doing the predicting cannot see the answer — so the model is built with the future taped over. The technical name is the causal mask: a word may attend to itself and everything before it, never after.
This is also why one training sentence teaches thousands of lessons at once: every position in “the cat sat on the mat” is simultaneously a quiz — after “the”, guess; after “cat”, guess; after “sat”, guess.
Then it does that dozens of times
One attention pass plus one thinking pass is a layer. Real models stack many layers — the small model behind this page's curriculum runs twelve, and each word's notebook holds 768 numbers. Early layers settle simple facts about the words; deeper layers handle the harder business of what the sentence means and what should come next.
The whole loop, in three questions
If you can answer these, you can explain how a language model reads — which is most of what a transformer is.
That is the reading loop: a word becomes a place, attention gathers its context, the notebooks deepen, and the last state points at the next word. What the model does with that final point — how it actually picks — is the next page.