Reading a model's mind

linear probes, in plain words · no code, no math · about 8 minutes · all data illustrative
BLOCK 1 · THE IDEA

Don't listen to what it says. Look at what it thinks.

Ask a model to process two sentences: “The sky is blue” and “The sky is green”. Before it produces a single word of output, its inside state for each sentence is different — a different pattern of activity.

On this page, every dot is one statement's inside state, simplified to a point on a plane. Teal dots are true statements, rose dots are false ones. Hover or tap a few dots to read them.

true statements false statements
Something is already visible: the two kinds of statements land in different regions — before the model says a word.
BLOCK 2 · THE TOOL

The test is one straight line

Here is the whole trick of a linear probe: draw a straight line between the two regions. Everything on one side counts as true, everything on the other as false. That line — a direction plus a threshold — is the entire probe.

Drag the slider and watch the score. If one angle separates the clouds cleanly, then the true/false difference is a direction in the model's internals — the finding that started this whole field.

How would anyone build such a line in the first place? From pairs: the same statement said truly (“I liked the movie”) and falsely (“I hated the movie”). Many pairs go in; the direction that separates them comes out.

BLOCK 3 · DEPTH

Where in the model does the difference live?

A model reads its input through many layers, one after another. You can train a probe on the inside state after each layer and ask: how separable are true and false here? The bars below show that score, layer by layer.

Click a layer.

The shape is the finding: barely above a coin flip in the early layers, sharpening through the middle, saturating late. The model doesn't get told what is true — the difference between processing a true and a false statement appears on its own, and grows as the information deepens.

BLOCK 4 · THE STAKES

From truth to lies

The same probe idea works on something less innocent. Train it on dialogues where the model is told to lie versus told to tell the truth — then run it on new conversations, and ask one question of each reply: do the internals look like the lying pattern?

Click each reply below. The meter is the probe's read. The words are yours to judge — that is the point.

A low read means the internals look straightforward. A high read means they match the pattern the probe learned from told-to-lie dialogues.

In the third reply, the words are confident and unremarkable. The inside is what breaks the tie — that is why safety researchers care about probes: they can check the model's state without waiting for the model's behavior to give the game away.

BLOCK 5 · THE CATCH

A straight line is a clue, not a conviction

Before this becomes magic, the honest caveat. A probe can separate the examples for the wrong reason — riding on some quirk of how the examples were written rather than the true/false difference itself. And a line that separates today's examples can fail on new ones. Probes are evidence about the internals, strong enough to act on and too weak to end the argument.

What does a cleanly separating line actually establish?
Why do safety researchers care about probes at all?

This page's workshop went further: causal interventions that push a model's state along the truth direction and watch its answers change, and probes that run on attention patterns rather than plain states. Same idea, sharper instruments.

TARA COMPANION · WEEK 4 · LINEAR PROBES · ALL DATA SYNTHETIC, ILLUSTRATIVE · AFTER MARKS & TEGMARK 2023 AND APOLLO RESEARCH 2025