ESSAY · PROJECT 0003
An eval is a threat model with a scoreboard
Entry 6 of the ARENA guide. Part 3.1, intro to evals. Route position: program week 5, tagged [core], first week of the evals-and-agents phase. Reading thread: Seredyński et al., AI agents in algorithmic electricity markets.
The mental model
Chapter 3 exists because the capstone needs tools, and this part exists because an
eval is more than a quiz with grades. The week's actual order of operations, straight
from the page: pick a property you worry a model might have, write the **threat
model that says when and how it would matter, turn that into a specification**
(otherwise the eval measures whatever is easy), and only then engineer the questions.
The replication exercise lands early for the same reason: ARENA has you replicate
alignment faking before you write a single question, so you feel what a real eval
result looks like from the inside.
The vocabulary shift is the point. Capability evals ask "can the model do X"; property
or control evals ask "does the model do X under conditions that matter for safety".
Same machinery, different claim. Every exercise this week feeds the capstone's threat
model, which the route schedules for this very week.
The picture
The exercises
Quoted from the pinned 3.1 page, with ARENA's own budgets:
- retry with exponential back-off (optional): 10–15 min
- replicate alignment faking: 25–45 min
- build threat models: 40–60 min
- design a specification: 45 min
- engineer the eval question prompt: 35–40 min
What it unlocks
Everything downstream is this part's pipeline wearing different clothes: 3.2 fills the
questions, 3.3 runs them, 3.4 points them at agents, 3.5 assumes the model is
adversarial. The capstone's own threat model is written this week with the tools
still warm.
Budget and receipts
ARENA's budgets sum to roughly 2.5–3.5 hours, inside the ~4-hour slot. The threat
model and specification exercises are where the hours should go; the optional backoff
can slip without guilt.
[TODO: receipts — hours vs budget, the week's threat-model claim, what broke (fill from the TARA vault when the week closes)]