How a language model reads

the transformer, rebuilt in plain words · no code, no math · about 8 minutes
BLOCK 1 · THE FIRST TRICK

A word becomes a place

A model cannot read letters. Every word it takes in is turned into a list of numbers — a point in a huge space. The trick that makes this useful: words that behave alike end up close together in that space.

Click a word on the map. The lines show its nearest neighbors — the words the model has learned to treat as its relatives.

Click any word to see its two nearest neighbors.
Two words sitting close in this space means…
BLOCK 2 · THE WORKSPACE

Every word carries a running notebook

As the sentence travels through the model, each word keeps one running list of numbers — its own notebook of what it means so far. Every layer reads the notebooks and adds notes. Nothing is ever erased; the model only ever adds.

Click each stage of the trip. (The real name for the notebook is the residual stream — one of only five terms this page will ask you to keep.)

→ → → →
Stage 1. The word "cat" becomes a point in the space from Block 1. Where it sits already says something; the next stages add to it.
BLOCK 3 · THE CORE IDEA

Attention: each word gathers its context

The first pass inside the model is attention. Each word looks back at the words before it and copies a little from each — the lines below, with thickness meaning how much gets copied.

Click the word “it”. Who does it belong to? The model settles that by copying hardest from one earlier word.

BLOCK 4 · THE RULE

Why a word may never look ahead

The toggle in Block 3 was not decoration. A language model is trained for one job: given everything so far, predict the next word. That job is only honest if the word doing the predicting cannot see the answer — so the model is built with the future taped over. The technical name is the causal mask: a word may attend to itself and everything before it, never after.

This is also why one training sentence teaches thousands of lessons at once: every position in “the cat sat on the mat” is simultaneously a quiz — after “the”, guess; after “cat”, guess; after “sat”, guess.

Why is the future taped over?
BLOCK 5 · DEPTH

Then it does that dozens of times

One attention pass plus one thinking pass is a layer. Real models stack many layers — the small model behind this page's curriculum runs twelve, and each word's notebook holds 768 numbers. Early layers settle simple facts about the words; deeper layers handle the harder business of what the sentence means and what should come next.

L 1–2Word identity and grammar — who is doing what to whom.
L 3–8Phrases bind into facts: "sat on the mat" stops being five words and becomes one situation.
L 9–12The notebooks are mostly finished products — ready to score what the next word should be.
Click a band of layers. (The split is illustrative — real models mix these jobs everywhere.)
BLOCK 6 · CHECK YOURSELF

The whole loop, in three questions

If you can answer these, you can explain how a language model reads — which is most of what a transformer is.

1. "Cat" and "dog" end up close together because…
2. In the sentence from Block 3, the word "it" works out what it refers to by…
3. Deeper layers matter because…

That is the reading loop: a word becomes a place, attention gathers its context, the notebooks deepen, and the last state points at the next word. What the model does with that final point — how it actually picks — is the next page.

TARA COMPANION · WEEK 2 · TRANSFORMERS FROM SCRATCH · ALL DATA HAND-SET, ILLUSTRATIVE · ARENA 3.0 PART 1.1