Concept register · Concept 42 of 64 · Theme: outcomes, not activity Reviewed 2026-09-01

assay  ·  concepts  ·  outcomes-not-activity

Reliability nines, not capability, as the binding axis

Reliability is a distinct capability axis from breadth, fluency and plasticity, and it has lagged behind them. That gap is why demos are easy and shipped agents fail — and each additional nine is bought by redesign, not by a better model.

established · assay: hard gates, nines unmeasured

3 independent sources  ·  sighted at DevCon London 2026 and the Agentic AI Summit 2026  ·  last reviewed 2026-09-01


§1What it is

A threshold below which automation is not automation

A workflow that is right most of the time does not remove human work; it re-prices it. Every wrong suggestion costs attention and burns the trust the whole system runs on, and interrupting someone’s flow to be wrong is expensive in a way a benchmark never shows. Near-total reliability is the line between automating a job and merely streamlining the human doing it.

Each nine is bought by redesign

99% works at small scale. 99.9% and then 99.999% each require a new system — new layers, new enforcement points, a different architecture. A better model alone never buys a nine. This is the observation from a decade of physical autonomy, and it reframes reliability work from incremental hardening into deliberate re-architecture.

The barriers are structural

Five of them, and none is incremental. Errors are invisible rather than loud. Control surfaces are soft — instructions, not enforcement. Metacognition is missing rather than weak, so a model cannot distinguish knowing something from a plausible guess. Much of what matters is not verifiable by RL. And the workarounds cost money and latency. Audits find real hallucination rates roughly 5x what customers believed they were living with.


§2Sightings

DevCon London 2026 · 1 sighting

Agentic AI Summit 2026 · 2 sightings

Also: Scale Cognition APT — specialist self-verifying models, with humans on one side and APIs on the other.


§3Where Assay stands

An instance of the workaround, not a solution to it

Assay’s desk machinery is precisely the external harness for hard control that this evidence names as the current state of the art — and it pays every cost attached to that approach: latency, expense, and fragility. That is the honest read. Assay is not ahead of this problem; it is an instance of the workaround (desk roles).

Two positions the evidence supports directly

Server-side enforcement is a redesign-for-the-nine, not a hardening chore. The rule that client-side guards are advisory and only server-side controls actually bind is exactly the instinct here — you buy a nine by adding an enforcement layer, not by making the existing layer try harder. Shipped, and the framing is worth keeping. The 5x hidden-error figure strengthens independent re-execution. The rule that a non-implementer re-runs the brief’s checks on merged main, rather than trusting worker-reported success, is the mechanism that catches errors an agent does not know it made (lifecycle). That audit result is the strongest external number supporting it.

No nines target on anything

There is no stated reliability bar for a desk loop, no measurement of how often a desk gets it wrong, and therefore no way to tell whether the loops are running at 99% or 99.9%. Moving them from mostly-working to unattended is a redesign problem, and Assay has no current estimate of which nine it is on. Hardening work is redesign in this sense, but it is not framed or measured against a reliability target. The 90% threshold is also the one-sentence rationale for the standing preference for pushing judgement into deterministic tooling and keeping agent loops on evidence rows rather than vibes: a probabilistic gate at 90% does not remove the human, it just moves where the human is interrupted.


§4Watch

  • A published nines figure for an agentic software workflow — the existing numbers come from physical autonomy, and the transfer to coding agents is asserted rather than measured.
  • Whether specialist self-verifying models actually move the accuracy-versus-cost frontier in third-party hands.
  • Any account of what redesign bought a specific nine in an agent system, as opposed to the general claim that redesign is required.