Concept register · Concept 52 of 64 · Theme: agent evals Reviewed 2026-09-01
assay · concepts · agent-evals
Capability-ceiling evals vs deployment-readiness evals
Public benchmarks are designed against saturation — their authors keep tasks hard so the leaderboard stays informative — which means a benchmark score measures a capability ceiling. Deciding whether to put an agent in front of real work is a different question, and needs a different exam: a per-use-case qualification with a stated risk envelope, an oversight policy, and an improvement path.
corroborated · assay: verify rows, three gaps
9 independent sources · sighted at the Agentic AI Summit 2026 · last reviewed 2026-09-01
§1What it is
Two different exams
The gap is quantitative and structural at once. Top text-to-SQL benchmarks sit at 60–70% because their designers avoid saturation, while enterprises buy at 99% and above; the leaderboard and the purchase decision are simply not measuring the same thing. The readiness exam also decomposes differently. An agentic evaluation is not an input/output pair but five gradeable parts — the task, the trajectory of reasoning and tool calls, the artifact produced, the world or context traversed, and the rubric or verifier.
Five primitives a capability benchmark never scores
Readiness adds axes the ceiling exam has no reason to carry: grounding precision and recall on citations; policy compliance; calibrated knowledge of what the agent does not know, with an explicit hand-to-a-human policy; safe state mutation, meaning recoverability, rollback and privacy; and the economics of the human oversight the deployment implies. The closing formula is not how intelligent the agent is, but what it is ready to do, under what constraints, at what risk, at what cost.
Build the exam where you will deploy
Two disciplines fall out. Shape the environment like the one you will deploy into — messy, incompletely documented, iterative — because a benchmark whose English maps directly onto the query clause tests nothing, and an ambiguous question can mark a better analyst’s richer answer wrong. And anchor the task set to economically real work rather than research convenience: today’s agent benchmarks cover programming and mathematics, roughly 7% of US employment.
§2Sightings
Agentic AI Summit 2026 · 11 sightings
#153The Exam Before Enterprise DeploymentYuan Emily Xue, Scale AI
#142Data Benchmarks: Where Everything’s Made UpGrace Tang, Hex
#144From Training to Evaluation: Open Recipes for Agentic AIChenguang Wang, UCSC / Scale AI
#158ScarfBench: Can Agents Migrate Enterprise Java?Rahul Krishna, IBM Research
#120Workshop: Future of Agent EvaluationBerkeley RDI and others
#001Enterprise AIAdarsh Hiremath, Mercor
#033Panel: Enterprise AILennox, HubSpot; Surapaneni, Google Cloud; Anita, Snowflake; Hiremath, Mercor; Sergey, Ema
#043Panel: Agentic AI in Finance & LegalChandhok, Circle; Shafiq, Wells Fargo
#030Opening Remarks, Day 2Dawn Song, Berkeley RDI
#038Panel: Frontier ResearchChi, Google; Socher, Recursive; Cubuk, Periodic Labs
#160Towards AI Co-ScientistsRose Yu, UC San Diego
Also: DS-Bench; Spider 2; DABS-Step; Shoreline; SWE-Atlas; Program Bench; OSWorld; Agent’s Last Exam; O*NET and the US Standard Occupational Classification; AgentExam; AgentBees; the Mercor Apex benchmark.
§3Where Assay stands
Verify rows are this craft institutionalized
A brief that cannot state a falsifiable check is a brief that has not been specified, and the review gate before authoring is the point at which that is caught — which is the readiness discipline written into a methodology rather than into a benchmark suite. The five-part anatomy maps cleanly onto what a brief already carries: verify rows are the task-plus-rubric slice, and evidence-cited verdicts are the trajectory-plus-artifact slice.
Three things the frontier does that Assay does not
Grounding precision and recall as a quality bar for evidence rows. An unverifiable citation fails precision; ignoring part of the brief fails recall. Today evidence rows are checked for presence and for lint, not scored on either axis. Specificity testing of verify rows. The ambiguous-question pathology — a better interpretation marked wrong — is exactly what an under-specified verify row does to a worker, and testing rows for specificity at authoring time is designed nowhere. Calibration and abstention as graded behaviors. Abstention is already first-class in the vocabulary — blocked-on-human and needs-decision are sanctioned exits, not failures — but nothing computes whether a desk’s confidence tracks its outcomes (desk roles).
Two findings that land as warnings
Migration asymmetry means an “X to Y” brief and a “Y to X” brief are different-difficulty work and should not be sized alike. And the human-anchored ceiling — a benchmark built from human labels cannot measure past human level — warns that human-authored verify rows cap what Assay can measure at all; environment-based verification through CI, canaries and production signals is the escape hatch, and it is the one the verify role already prefers. Labor-anchored evaluation is the standard any future outcome-metrics work would need to aim at if it moved from counting activity to measuring economically real output. That is a stated ambition, not a shipped capability.
§4Watch
- A first-hand account of a readiness gate actually blocking a deployment — the primitives are well specified, but the evidence is still framework-and-benchmark work rather than deployment decisions. That would move this to established.
- Whether readiness primitives converge on a shared vocabulary across vendors, or stay per-vendor.
- Whether calibration gets measured in practice — a published agent calibration curve against outcomes, rather than calibration named as a desirable property.