Everything a chatbot does starts as this game
Given everything so far, guess the next word. That is the whole task a language model trains on — billions of times. Fill in the blank yourself first, then see the model's own betting sheet.
Scores become shares of 100%
Before the percentages there are raw scores — the model produces one for every word it knows, and they can be any number at all. A fixed recipe (softmax, the only name worth keeping) turns scores into shares of 100%: a word's share grows exponentially with its score, and every word's gain is another word's loss.
Pick a word, then drag its score up and down. Watch the others.
Temperature: how spicy the model feels like being
Percentages are not destiny. At the final step, the shares get reshaped by a dial called temperature before one word is drawn. Low temperature sharpens the favorites; high temperature fattens the long shots.
Guess, miss, nudge
Where did the scores come from? From practice. Training is a loop: the model guesses, the truth makes the miss measurable, and every one of its billions of knobs gets nudged a little in the direction that would have made the guess better. Then again. And again.
Below, a toy model with five words learns from exactly three sentences: “the cat sat”, “the cat ran”, “the dog sat”. Its only job: after “the”, predict the next word. It starts knowing nothing.
From this page to the whole system
And the bridge to the next page: the scores came from the model's inside state — the very thing week 4's straight-line test reads. A probe does not listen to the words; it reads the state that produced them.