ESSAY · PROJECT 0002
You can read a model's internals with one straight line
This week's technical workshop was the one the safety people care about most: linear probes. Ask a model to process a statement, then look not at its answer but at its internal state. True statements and false statements land in different regions. A straight line between those regions is an entire lie detector. The page lets you draw the line yourself, then points it at deception.
Where this comes from
- The workshop notebook: ARENA 3.0, part 1.3.1: Linear Probes — truth representations, probe training, causal interventions, deception probes, attention probes.
- The discovery: The Geometry of Truth (Marks & Tegmark, 2023), linear representations of truth that generalize across datasets.
- The stakes: deception probes (Apollo Research, 2025), probes trained on simple contrastive data catching role-played deception in realistic scenarios; and high-stakes interactions (NeurIPS 2025).
- Every dot, score, and dialogue on the page is synthetic and illustrative: the shape of the findings, not measurements from a real model.