Key terms brief register desk statusgen tombstone

assay · how it runs · the pipeline, measured on itself

How it runs

These are the instruments for our own pipeline. The headline is the AssayScore, one 0–100 number for how this project turns briefs (units of work, sized and gated before dispatch) into delivered work, computed from its own git and artifact record rather than from a survey. It comes with its four sub-scores, the raw inputs behind each, and the reference bands they were normalized against, so the composite can be checked against its parts. Everything the tooling could not compute is published too, as could-not-check, never as a zero.

Snapshot taken2026-09-06T13:25:59Z
Source commitce9e696
Windownot stated in this snapshot
AssayScore16.7 / 100 (measured)

The AssayScore

Composite, 0–100

16.7/100

measured

Computed over all four dimensions, scaled against this project's own trailing-90-day history. Self-relative — not a comparison with anybody else.

The four sub-scores, each 0–100
Speed57.7

brief lead time, inverse-scaled against the trailing-90-day band

Flow0.9

flow efficiency, a bounded ratio scaled to 0–100

Quality86.7

first-pass yield, a bounded ratio scaled to 0–100

Value17.4

weighted throughput per hour of human decision time, scaled against the trailing-90-day band

Each bar is that dimension's own 0–100 sub-score, and the value is printed beside it. A dimension that could not be computed shows an empty track and the words could-not-check. It is never drawn as a zero-length bar, because a bar at zero and a bar that was never measured look identical and mean opposite things.

How the score is built

It is the geometric mean of the four sub-scores, (Speed · Flow · Quality · Value) to the power of a quarter, not an average. A geometric mean penalises imbalance on purpose: one weak dimension pulls the composite down and cannot be offset by a strong one. That is what makes the number hard to game, and it also means a low composite usually names one weak dimension rather than a broad failure. Read the four bars, not the headline.

Two of the four are ratios; two are normalized against this project's own history. Flow is flow efficiency and Quality is first-pass yield. Both already sit in 0–1 and map straight onto 0–100 with no baseline involved. Speed (brief lead time, where lower is better) and Value (weighted throughput per hour of human decision time) are unbounded, so each is scaled against a reference band taken from its own observations over the trailing ninety days, the 10th and 90th percentile, so one outlier cannot move the band. The exact band, its date range, and the number of observations behind it are published below and in the snapshot.

The raw inputs behind each sub-score, and the band it was scaled against
DimensionSub-scoreRaw input, as the emitter reported itReference band
Speed
speed
57.7median lead time 17.7 days · n = 336p10 3.7 to p90 36.8, over n = 384
Flow
flow
0.9flow efficiency 0.9%no band — bounded ratio
Quality
quality
86.7first-pass yield 86.7% · 13 first-pass · n = 15no band — bounded ratio
Value
value
17.4920 effort points · 6621.4 human-decision hours · points per hour 0.1389 · denominator = decision-queue dwell; review-latency + gate-touch terms are future substratep10 not reported to p90 0.8, over n = 6

Baseline window for the normalized dimensions: 2026-06-08 to 2026-09-06 — the percentile bands above are taken from this project's own observations across that range, and from nowhere else

How to read these numbers

These are diagnostic, not a target. The emitter says so in its own output, and it is repeated here because that is the basis on which these numbers are published at all. A delivery measure that becomes a target stops being a good measure. If one of these numbers starts driving behaviour, that is a finding to raise, not a score to improve.

They are derived from agent-authored artifacts with consistency linting, not measured from ground truth. The lifecycle states these metrics are computed over are written by the same automated sessions that do the work. Consistency linting makes contradictions machine-visible; it does not make the underlying claims true, and a state recorded by whoever did the work is not independent evidence that the work was done. Read every figure here as the system's account of itself.

A count is not a claim. The methodology overview states that no productivity multiplier, output-per-agent figure, or tier comparison gets published, because there is no baseline against which such a figure could be recomputed. That still holds, and this page is not an exception to it. The composite at the top is scaled against this project's own history and says nothing about anyone else's; the raw counts of our own activity are recomputable by anyone holding the same command, the same window, and the same commit. None of it asserts that the work was faster, cheaper, or better than any alternative, and no figure here should be read as making that assertion.

This is one repository, over one window, at one instant. It is not a fleet roll-up, not a benchmark, and not a comparison against anybody. Counts move while they are being taken, because the corpus changes during the measurement, so every number on this page is stamped with the instant and the source commit it was taken at, and means nothing without them.

One reference, and it is us. There are no external deployments yet. Every operational number here comes from our own fleet, the toolkit's own development, and it is labelled as such. We run our own company on it: one human, an agent fleet. It is the only reference that exists today, and you can read all of it. The same pipeline seen as a live board is on the board.

The flow metrics underneath the score

These are the per-brief flow measurements the score is built out of, plus the ones published beside it as context. When the composite moves, this is the table that says which way and why. Each is in the same three states as everything else on this page, measured, partial, or could-not-check, and a metric that could not be computed keeps its row and prints those words rather than showing a zero.

Brief-flow metrics, each in one of three states
MetricValueStateWhat it counts
Throughput
throughput
976 effort points · 344 briefs · 264 of unknown sizemeasuredbriefs delivered in the window, and their weighted effort points
Lead time by size
leadtime
S: 26.7 d median, 34.8 d p85, n = 103 · M: 17.7 d median, 36.8 d p85, n = 203 · L: 14.7 d median, 29.1 d p85, n = 30measuredmedian and 85th-percentile days from authored to done, split by brief size
Flow efficiency
flow_efficiency
0.9% of elapsed time worked · 12 working spans against 1232 waitsmeasuredthe share of elapsed time a brief was actually being worked rather than waiting
First-pass yield
first_pass_yield
86.7% · 13 of 15 briefs · 45 excluded as unlinkedmeasuredthe share of briefs that reached done without a rework round
Review rework
review_rework
0.40 review rounds per brief · n = 40measuredmean review rounds per brief, and how the rounds are distributed
Decision latency
decision_latency
p50 74.6 hours · p90 429.7 hours · 22 open now · oldest waiting 523.2 hoursmeasuredhow long an open decision waits on a human, and how many are open now
Net flow and stalls
stall
could-not-checkcould-not-checkthe per-stream net-flow / stall feed was not probed for this snapshot

The five delivery metrics

The commodity delivery family is published alongside, from its own feed. It is not folded into the composite. Each metric is in exactly one of three states. measured means the emitter computed it from a complete input. partial means it computed it from an input it itself declares incomplete, and the caveat travels with the number. could-not-check means it could not compute it at all. A could-not-check metric is never shown as zero and never dropped from the table: a zero for a probe that did not run is a false claim, and a dropped row hides that this page's coverage is partial.

The five delivery metrics, each in one of three states
MetricValueStateWhat the emitter reports about it
Deployment frequency
deployment_frequency
could-not-checkthe feed did not carry this metric at all
Change lead time
change_lead_time
could-not-checkthe feed did not carry this metric at all
Failed-deploy recovery time
failed_deploy_recovery_time
could-not-checkthe feed did not carry this metric at all
Change failure rate
change_failure_rate
could-not-checkthe feed did not carry this metric at all
Rework rate
rework_rate
could-not-checkthe feed did not carry this metric at all
Coverage

Coverage of this table right now: 0 of 5 measured, 0 partial, 5 could-not-check — the commodity delivery (DORA) and velocity metrics were split out of this generator to a dedicated open-source metrics platform (Apache DevLake) and are served from there. They are not yet wired back into this snapshot, so the DORA series is UNREAD here — never a zero that reads as measured.. What cannot be computed is not a broken measurement. It is a measurement with no independent input yet. Publishing those as blanks would imply the system had nothing to report; publishing them as zero would imply it had nothing to fix. The last column is the emitter's own wording, kept verbatim rather than paraphrased, because a paraphrase would put a second author between you and what was reported. It names the desks (the standing roles that run the pipeline) that would supply the missing inputs; those roles are described in the methodology overview.

Trend over time

A trend view rolls the recorded lifecycle transitions up into a time series. Its state for this snapshot: could-not-check

the trend feed was not probed for this snapshot

Where each number comes from, and where it stops

Some inputs are list calls with a hard result cap. A count taken from a capped list call becomes a silent undercount the moment the cap binds: the call still reports success, and nothing downstream can tell the difference. The emitter's output does not state its own caps, so they are recorded explicitly when the snapshot is refreshed and published here alongside the numbers they bound.

Each probe's result cap, and what happens when it binds
ProbeBoundsResult capBehaviour at the cap
This snapshot states no probe caps, because the capped list calls behind the delivery table were not run for it.

Raw activity

Commits per day, merges per day, and pull requests per day used to be the headline of this page. They are not any more, and they are not gone either. They are here, one fold down, because deleting a number you have stopped leading with is a different and worse act than demoting it.

Show raw activity: commits, merges, and pull requests per day
Raw activity counters, per day
CounterPer dayStateWindow
Commits140.07measuredtrailing 28d (2026-08-09 to 2026-09-06)
Merges36.11measuredtrailing 28d (2026-08-09 to 2026-09-06)
Pull requests33.86measuredtrailing 28d (2026-08-09 to 2026-09-06)

raw activity is shown only behind a fold; the headline is the AssayScore. Commits, merges and PRs count motion, not delivered value.

A repository can produce a great many commits per day while shipping less than it did last month, and any of these three numbers can be raised on purpose by anyone who decides to raise it, which is why none of them is the metric of record here. They stay published because they are cheap to recompute and easy to check, and because a page that quietly dropped its old headline the day it became inconvenient would be making a claim of its own.

What this page does not cover

  • One repository only. The toolkit's own development repository. No other repository contributes a single number here, and nothing on this page should be read as a fleet-wide or organisation-wide figure.
  • One window only, the dates stamped at the top. Nothing before the window start is counted; a longer or shorter window would produce different numbers, and no window here was chosen to flatter a result.
  • No cross-organisation comparison. The composite is scaled against this project's own trailing history, so it is not comparable with a score computed against a different project's history. Publishing it as though it were would be the exact overclaim this page exists to avoid.
  • Deployments are not production deployments. The delivery throughput figure counts commits and merges in the repository, which is what the emitter can see. It is not a count of releases reaching users, and should not be read as one.
  • Change failure is a partial signal. It is computed from bug-labelled issues against merged changes. It does not include failures that were never filed as a bug-labelled issue, so it is a floor rather than a rate.
  • Not everything in the feed is drawn here. The published snapshot carries additional diagnostic blocks, the constraint stage, the intake front-door backlog, and the human-decision queue, that this page does not yet render. They are in the feed and readable there, and the board page draws them; saying so is cheaper than pretending this page is the whole of it.
  • No per-person anything. These are system-level aggregates. Nothing here is attributed to an individual, and the toolkit's own framing forbids using any of it as an individual scorecard.

The source this page was rendered from

Every number above is derived from one declared source, and nothing on this page is typed by hand. That source is published as metrics.assay.guide/metrics.json: the emitter's raw output, verbatim, plus the snapshot's commit, window, and probe caps. This page carries no numbers of its own. It is a static shell that fetches that feed live and renders it in your browser, so a fresh snapshot appears without redeploying the site. If the live feed cannot be reached it falls back to the copy committed next to the page.

The snapshot is regenerated on a schedule by the toolkit's own emitter and uploaded to the feed; the site itself is not rebuilt to move the numbers. A check re-renders the page shell and requires it to be byte-identical to what is published, which proves no number was hand-typed into the markup, and separately requires the published snapshot to be well-formed. This is the same stance the toolkit takes everywhere else: one writer, one source, and a script that can tell when the two have drifted.

The emitter's own framing, verbatim: the commodity delivery (DORA) and velocity metrics were split out of this generator to a dedicated open-source metrics platform (Apache DevLake) and are served from there. They are not yet wired back into this snapshot, so the DORA series is UNREAD here — never a zero that reads as measured.