Concept register · Theme 09 of 14 4 concepts · 34 talks

assay  ·  concepts  ·  agent-evals

Agent evals

Evaluation is moving off the pre-launch checkpoint: scored continuously on live traffic, over the trajectory rather than only the artifact, by verifiers that are themselves audited — and capability scores are finally distinguished from deployment decisions.

assay: readiness exam shipped · live evals open

Production outcome evals — evaluation as a standing process, not a gate — 10 sources, establishedTrajectory grading — scoring the path, not just the artifact — 10 sources, establishedBenchmark integrity and run variance — QC for the measuring stick — 9 sources, establishedCapability-ceiling evals vs deployment-readiness evals — 9 sources, corroborated
4 concepts · 38 independent sources · 3 established

§1What it is

A standing process, not a gate

Production outcome evals run continuously against live traffic and are measured by what changed because of them, not by a headline pass rate. Three shapes recur: always-on outcome evals, shadow mode (the agent proposes, the human does, the diff feeds learning), and causal in-the-wild evaluation that defines its estimand before it counts.

Score the path, not just the artifact

Trajectory grading treats the reasoning path and the tool-call sequence as a first-class object alongside the artifact, because outcome-only grading passes an agent that arrived by a forbidden or wasteful route — the trajectory is where the bugs live.

QC the measuring stick

Benchmark numbers routinely fail to mean what they appear to: saturation and contamination, verifiers with their own error rates, agents attacking the eval harness, and run variance a single number hides. And a benchmark score measures a capability ceiling — deciding to deploy is a different exam, a per-use-case qualification with a risk envelope and an oversight policy.


§2The concepts in this theme

Each concept has its own page in the concept register — with sightings from every event we review, and where Assay stands on each.


§3How Assay implements this

The readiness exam is shipped

A brief that cannot state a falsifiable check is a brief that has not been specified, and the review gate before authoring is where that is caught — the readiness discipline written into a methodology rather than a benchmark suite. Verify rows are the task-plus-rubric slice; evidence-cited verdicts are the trajectory-plus-artifact slice. Abstention is already first-class: blocked-on-human and needs-decision are sanctioned exits, not failures.

Pre-merge, and once

Verify rows run before a merge and run once. That is the honest position: nothing always-on exists — no post-merge standing evaluation, no judge agent scoring live outcomes, no re-run of a brief’s verify rows after the world moves under them. And nothing ties a landed brief back to whether a higher-level outcome moved; that absence is precisely why the recurring reports count activity.

The named gaps

No trace-level metric of any kind exists — “tokens per verified brief” is the concrete upgrade over activity counting. The verifiers have no audit of their own: a verify row that passes on a sabotaged submission is a broken row, and sabotage runs as a routine audit of row quality is the single cheapest steal in this theme. Verify rows also stay closed-form and adversary-resistant rather than judged by a model — a verify row is a tiny verifier, and verifiers are the thing being attacked.


§4Talks that cover this theme