sammatuba

ESSAY · PROJECT 0003

Reasoning models say more than they do, so measure the gap

Dataset the ARENA 3.0 curriculum (TARA vault submodule, pinned 81533b2) + the TARA Learning Plan v2 · Charts 0 · Written SEP 2026

Entry 13 of the ARENA guide. Part 4.3 §1–2, interpreting reasoning models. Route position: program week 12, tagged [light], API-only, the week of the W47 project-proposal gate. Reading thread: Meek et al., measuring chain-of-thought monitorability.

The mental model

Reasoning models produce a chain of thought, and safety work leans on it: if the

model says why it did something, a monitor can read that. The uncomfortable question

this part formalizes: is the chain of thought faithful to the computation that

produced the answer, or is it a plausible story told after the fact? "Interpreting

reasoning models" turns that worry into measurement.

The exercises build the measurement in layers. Categorize reasoning sentences

heuristically, then with an autorater. Then the core idea, importance: how much

does the final answer depend on a given piece of reasoning? Forced-answer importance

tests it directly (truncate the reasoning, force an answer, see if it changes);

resampling and counterfactual importance statisticalize it. The payoff exercise

replicates a published figure on how sentence categories in the chain relate to

outcomes. The habit being installed: never quote a model's self-report as evidence

without a faithfulness number attached.

The exercises

Quoted from the pinned 4.3 page, with ARENA's own budgets:

What it unlocks

This closes the core route. The proposal gate the same week demands an accepted

research intent and a design the mentor can interrogate, and CoT-monitorability is

exactly the kind of claim the capstone will make about its own evaluation. After the

gate, the ARENA slot becomes the build: weeks 13–14 belong to the capstone and demo

day, and this guide's remaining pages are written by the work.

Budget and receipts

ARENA's budgets sum to roughly 1–1.5 hours, deliberately light: the week's real

deliverable is the proposal. The route tags the part [light] and API-only; no GPU

appears.

[TODO: receipts — hours vs budget, faithfulness measurement results, the proposal's fate at the gate (fill from the TARA vault when the week closes)]

PREVIOUS Week 12 is a fork, and either road is one session long · GUIDE HOME