Don't listen to what it says. Look at what it thinks.
Ask a model to process two sentences: “The sky is blue” and “The sky is green”. Before it produces a single word of output, its inside state for each sentence is different — a different pattern of activity.
On this page, every dot is one statement's inside state, simplified to a point on a plane. Teal dots are true statements, rose dots are false ones. Hover or tap a few dots to read them.
The test is one straight line
Here is the whole trick of a linear probe: draw a straight line between the two regions. Everything on one side counts as true, everything on the other as false. That line — a direction plus a threshold — is the entire probe.
Drag the slider and watch the score. If one angle separates the clouds cleanly, then the true/false difference is a direction in the model's internals — the finding that started this whole field.
How would anyone build such a line in the first place? From pairs: the same statement said truly (“I liked the movie”) and falsely (“I hated the movie”). Many pairs go in; the direction that separates them comes out.
Where in the model does the difference live?
A model reads its input through many layers, one after another. You can train a probe on the inside state after each layer and ask: how separable are true and false here? The bars below show that score, layer by layer.
The shape is the finding: barely above a coin flip in the early layers, sharpening through the middle, saturating late. The model doesn't get told what is true — the difference between processing a true and a false statement appears on its own, and grows as the information deepens.
From truth to lies
The same probe idea works on something less innocent. Train it on dialogues where the model is told to lie versus told to tell the truth — then run it on new conversations, and ask one question of each reply: do the internals look like the lying pattern?
Click each reply below. The meter is the probe's read. The words are yours to judge — that is the point.
In the third reply, the words are confident and unremarkable. The inside is what breaks the tie — that is why safety researchers care about probes: they can check the model's state without waiting for the model's behavior to give the game away.
A straight line is a clue, not a conviction
Before this becomes magic, the honest caveat. A probe can separate the examples for the wrong reason — riding on some quirk of how the examples were written rather than the true/false difference itself. And a line that separates today's examples can fail on new ones. Probes are evidence about the internals, strong enough to act on and too weak to end the argument.
This page's workshop went further: causal interventions that push a model's state along the truth direction and watch its answers change, and probes that run on attention patterns rather than plain states. Same idea, sharper instruments.