ESSAY · PROJECT 0003
Reasoning models say more than they do, so measure the gap
Entry 13 of the ARENA guide. Part 4.3 §1–2, interpreting reasoning models. Route position: program week 12, tagged [light], API-only, the week of the W47 project-proposal gate. Reading thread: Meek et al., measuring chain-of-thought monitorability.
The mental model
Reasoning models produce a chain of thought, and safety work leans on it: if the
model says why it did something, a monitor can read that. The uncomfortable question
this part formalizes: is the chain of thought faithful to the computation that
produced the answer, or is it a plausible story told after the fact? "Interpreting
reasoning models" turns that worry into measurement.
The exercises build the measurement in layers. Categorize reasoning sentences
heuristically, then with an autorater. Then the core idea, importance: how much
does the final answer depend on a given piece of reasoning? Forced-answer importance
tests it directly (truncate the reasoning, force an answer, see if it changes);
resampling and counterfactual importance statisticalize it. The payoff exercise
replicates a published figure on how sentence categories in the chain relate to
outcomes. The habit being installed: never quote a model's self-report as evidence
without a faithfulness number attached.
The exercises
Quoted from the pinned 4.3 page, with ARENA's own budgets:
- heuristic-based categorization: 10–15 min
- implement an autorater: 10–15 min
- calculate forced answer importance: 10–15 min
- compare resampling importance: 5–10 min
- compute counterfactual importance: 15–20 min
- replicate Figure 3b (sentence category effect): 10–15 min
What it unlocks
This closes the core route. The proposal gate the same week demands an accepted
research intent and a design the mentor can interrogate, and CoT-monitorability is
exactly the kind of claim the capstone will make about its own evaluation. After the
gate, the ARENA slot becomes the build: weeks 13–14 belong to the capstone and demo
day, and this guide's remaining pages are written by the work.
Budget and receipts
ARENA's budgets sum to roughly 1–1.5 hours, deliberately light: the week's real
deliverable is the proposal. The route tags the part [light] and API-only; no GPU
appears.
[TODO: receipts — hours vs budget, faithfulness measurement results, the proposal's fate at the gate (fill from the TARA vault when the week closes)]