ESSAY · PROJECT 0002
The model doesn't know the next word — it bets on it
Week 2 ended with a model whose final state "points at" the next word. This page is what happens at that final step and after: raw scores become shares of 100%, a dial called temperature reshapes the shares before one word is drawn, and the scores themselves came from a loop of guess, miss, nudge. You can play the whole training loop on the page; it is small enough to watch think.
Where this comes from
- The curriculum: ARENA 3.0, whose scoring step closes part 1.1's transformer, and the training loop is part 0.3's subject.
- For a video companion: Andrej Karpathy, "Intro to Large Language Models" , the betting-sheet framing, at full scale.
- The six candidate scores are hand-set; the toy training loop on the page is exact arithmetic on three sentences, not an illustration of arithmetic.
- Next in the series: the strangest fact of all: a straight line through the model's insides can tell true statements from false ones.