Concept register · Theme 09 of 14 4 concepts · 34 talks
assay · concepts · agent-evals
Agent evals
Evaluation is moving off the pre-launch checkpoint: scored continuously on live traffic, over the trajectory rather than only the artifact, by verifiers that are themselves audited — and capability scores are finally distinguished from deployment decisions.
assay: readiness exam shipped · live evals open
§1What it is
A standing process, not a gate
Production outcome evals run continuously against live traffic and are measured by what changed because of them, not by a headline pass rate. Three shapes recur: always-on outcome evals, shadow mode (the agent proposes, the human does, the diff feeds learning), and causal in-the-wild evaluation that defines its estimand before it counts.
Score the path, not just the artifact
Trajectory grading treats the reasoning path and the tool-call sequence as a first-class object alongside the artifact, because outcome-only grading passes an agent that arrived by a forbidden or wasteful route — the trajectory is where the bugs live.
QC the measuring stick
Benchmark numbers routinely fail to mean what they appear to: saturation and contamination, verifiers with their own error rates, agents attacking the eval harness, and run variance a single number hides. And a benchmark score measures a capability ceiling — deciding to deploy is a different exam, a per-use-case qualification with a risk envelope and an oversight policy.
§2The concepts in this theme
Each concept has its own page in the concept register — with sightings from every event we review, and where Assay stands on each.
- Production outcome evals — evaluation as a standing process, not a gate established 10
- Trajectory grading — scoring the path, not just the artifact established 10
- Benchmark integrity and run variance — QC for the measuring stick established 9
- Capability-ceiling evals vs deployment-readiness evals corroborated 9
§3How Assay implements this
The readiness exam is shipped
A brief that cannot state a falsifiable check is a brief that has not been specified, and the review gate before authoring is where that is caught — the readiness discipline written into a methodology rather than a benchmark suite. Verify rows are the task-plus-rubric slice; evidence-cited verdicts are the trajectory-plus-artifact slice. Abstention is already first-class: blocked-on-human and needs-decision are sanctioned exits, not failures.
Pre-merge, and once
Verify rows run before a merge and run once. That is the honest position: nothing always-on exists — no post-merge standing evaluation, no judge agent scoring live outcomes, no re-run of a brief’s verify rows after the world moves under them. And nothing ties a landed brief back to whether a higher-level outcome moved; that absence is precisely why the recurring reports count activity.
The named gaps
No trace-level metric of any kind exists — “tokens per verified brief” is the concrete upgrade over activity counting. The verifiers have no audit of their own: a verify row that passes on a sabotaged submission is a broken row, and sabotage runs as a routine audit of row quality is the single cheapest steal in this theme. Verify rows also stay closed-form and adversary-resistant rather than judged by a model — a verify row is a tiny verifier, and verifiers are the thing being attacked.
§4Talks that cover this theme
#043Panel: Agentic AI in Finance & LegalChandhok, Circle; Shafiq, Wells Fargo
#110Agent Arena: Causal Evaluations of Agents in the Real WorldAnastasios N Angelopoulos, Arena
#144From Training to Evaluation: Open Recipes for Agentic AIChenguang Wang, UCSC / Scale AI
#001Enterprise AIAdarsh Hiremath, Mercor
#108The Art & Science of Benchmarking AgentsVincent Sunn Chen, Snorkel
#142Data Benchmarks: Where Everything’s Made UpGrace Tang, Hex
#153The Exam Before Enterprise DeploymentYuan Emily Xue, Scale AI
#120Workshop: Future of Agent EvaluationBerkeley RDI and others
8 of 34 talks shown — the ones that reach the most concepts in this theme. Every sighting, per talk, is on the concept pages above.