sammatuba

ESSAY · PROJECT 0003

An eval is a threat model with a scoreboard

Dataset the ARENA 3.0 curriculum (TARA vault submodule, pinned 81533b2) + the TARA Learning Plan v2 · Charts 1 · Written SEP 2026

Entry 6 of the ARENA guide. Part 3.1, intro to evals. Route position: program week 5, tagged [core], first week of the evals-and-agents phase. Reading thread: Seredyński et al., AI agents in algorithmic electricity markets.

The mental model

Chapter 3 exists because the capstone needs tools, and this part exists because an

eval is more than a quiz with grades. The week's actual order of operations, straight

from the page: pick a property you worry a model might have, write the **threat

model that says when and how it would matter, turn that into a specification**

(otherwise the eval measures whatever is easy), and only then engineer the questions.

The replication exercise lands early for the same reason: ARENA has you replicate

alignment faking before you write a single question, so you feel what a real eval

result looks like from the inside.

The vocabulary shift is the point. Capability evals ask "can the model do X"; property

or control evals ask "does the model do X under conditions that matter for safety".

Same machinery, different claim. Every exercise this week feeds the capstone's threat

model, which the route schedules for this very week.

The picture

The exercises

Quoted from the pinned 3.1 page, with ARENA's own budgets:

What it unlocks

Everything downstream is this part's pipeline wearing different clothes: 3.2 fills the

questions, 3.3 runs them, 3.4 points them at agents, 3.5 assumes the model is

adversarial. The capstone's own threat model is written this week with the tools

still warm.

Budget and receipts

ARENA's budgets sum to roughly 2.5–3.5 hours, inside the ~4-hour slot. The threat

model and specification exercises are where the hours should go; the optional backoff

can slip without guilt.

[TODO: receipts — hours vs budget, the week's threat-model claim, what broke (fill from the TARA vault when the week closes)]

PREVIOUS Mech interp reads the activations, not the outputs · NEXT Eval datasets are generated, not found · GUIDE HOME