How a model picks the next word

softmax, temperature and training, in plain words · no code, no math · about 8 minutes
BLOCK 1 · THE GAME

Everything a chatbot does starts as this game

Given everything so far, guess the next word. That is the whole task a language model trains on — billions of times. Fill in the blank yourself first, then see the model's own betting sheet.

The cat sat on the …
These are the model's probabilities for its six strongest candidates. Click one to complete the sentence.
BLOCK 2 · THE RECIPE

Scores become shares of 100%

Before the percentages there are raw scores — the model produces one for every word it knows, and they can be any number at all. A fixed recipe (softmax, the only name worth keeping) turns scores into shares of 100%: a word's share grows exponentially with its score, and every word's gain is another word's loss.

Pick a word, then drag its score up and down. Watch the others.

The shares always sum to 100%. "Mat" at 2.1 owns ~40%; drag it to −2 and watch the whole sheet redistribute — nothing is added, it is all competition.
BLOCK 3 · THE DIAL

Temperature: how spicy the model feels like being

Percentages are not destiny. At the final step, the shares get reshaped by a dial called temperature before one word is drawn. Low temperature sharpens the favorites; high temperature fattens the long shots.

The draw uses a seeded shuffle — deterministic, but you would not bet on it.
At 0.2 the model is a creature of habit; at 2.0 it reaches for "piano". This dial is why the same chatbot gives different answers to the same question.
BLOCK 4 · THE LEARNING

Guess, miss, nudge

Where did the scores come from? From practice. Training is a loop: the model guesses, the truth makes the miss measurable, and every one of its billions of knobs gets nudged a little in the direction that would have made the guess better. Then again. And again.

Below, a toy model with five words learns from exactly three sentences: “the cat sat”, “the cat ran”, “the dog sat”. Its only job: after “the”, predict the next word. It starts knowing nothing.

AVERAGE SURPRISE (THE "LOSS")
Press TRAIN STEP a few times. The bars walk toward the truth hidden in the three sentences — two-thirds "cat", one-third "dog" — and the surprise falls. Real training is exactly this loop, with billions of examples and billions of knobs.
BLOCK 5 · CHECK YOURSELF

From this page to the whole system

Where do the percentages come from?
Turning the temperature up…
One training step does what?

And the bridge to the next page: the scores came from the model's inside state — the very thing week 4's straight-line test reads. A probe does not listen to the words; it reads the state that produced them.

TARA COMPANION · WEEK 3 · SOFTMAX, TEMPERATURE, TRAINING · HAND-SET DATA, TOY LOOP IS EXACT · ARENA 3.0 PART 0.3/1.1