assay · how it runs · the pipeline, measured on itself
How it runs
These are the instruments for our own pipeline. The headline is the AssayScore, one 0–100 number for how this project turns briefs (units of work, sized and gated before dispatch) into delivered work, computed from its own git and artifact record rather than from a survey. It comes with its four sub-scores, the raw inputs behind each, and the reference bands they were normalized against, so the composite can be checked against its parts. Everything the tooling could not compute is published too, as could-not-check, never as a zero.
ce9e696The AssayScore
16.7/100
measured
Computed over all four dimensions, scaled against this project's own trailing-90-day history. Self-relative — not a comparison with anybody else.
brief lead time, inverse-scaled against the trailing-90-day band
flow efficiency, a bounded ratio scaled to 0–100
first-pass yield, a bounded ratio scaled to 0–100
weighted throughput per hour of human decision time, scaled against the trailing-90-day band
Each bar is that dimension's own 0–100 sub-score, and the value is printed beside it. A dimension that could not be computed shows an empty track and the words could-not-check. It is never drawn as a zero-length bar, because a bar at zero and a bar that was never measured look identical and mean opposite things.
How the score is built
It is the geometric mean of the four sub-scores, (Speed · Flow · Quality · Value) to the power of a quarter, not an average. A geometric mean penalises imbalance on purpose: one weak dimension pulls the composite down and cannot be offset by a strong one. That is what makes the number hard to game, and it also means a low composite usually names one weak dimension rather than a broad failure. Read the four bars, not the headline.
Two of the four are ratios; two are normalized against this project's own history. Flow is flow efficiency and Quality is first-pass yield. Both already sit in 0–1 and map straight onto 0–100 with no baseline involved. Speed (brief lead time, where lower is better) and Value (weighted throughput per hour of human decision time) are unbounded, so each is scaled against a reference band taken from its own observations over the trailing ninety days, the 10th and 90th percentile, so one outlier cannot move the band. The exact band, its date range, and the number of observations behind it are published below and in the snapshot.
| Dimension | Sub-score | Raw input, as the emitter reported it | Reference band |
|---|---|---|---|
| Speed speed | 57.7 | median lead time 17.7 days · n = 336 | p10 3.7 to p90 36.8, over n = 384 |
| Flow flow | 0.9 | flow efficiency 0.9% | no band — bounded ratio |
| Quality quality | 86.7 | first-pass yield 86.7% · 13 first-pass · n = 15 | no band — bounded ratio |
| Value value | 17.4 | 920 effort points · 6621.4 human-decision hours · points per hour 0.1389 · denominator = decision-queue dwell; review-latency + gate-touch terms are future substrate | p10 not reported to p90 0.8, over n = 6 |
Baseline window for the normalized dimensions: 2026-06-08 to 2026-09-06 — the percentile bands above are taken from this project's own observations across that range, and from nowhere else
How to read these numbers
These are diagnostic, not a target. The emitter says so in its own output, and it is repeated here because that is the basis on which these numbers are published at all. A delivery measure that becomes a target stops being a good measure. If one of these numbers starts driving behaviour, that is a finding to raise, not a score to improve.
They are derived from agent-authored artifacts with consistency linting, not measured from ground truth. The lifecycle states these metrics are computed over are written by the same automated sessions that do the work. Consistency linting makes contradictions machine-visible; it does not make the underlying claims true, and a state recorded by whoever did the work is not independent evidence that the work was done. Read every figure here as the system's account of itself.
A count is not a claim. The methodology overview states that no productivity multiplier, output-per-agent figure, or tier comparison gets published, because there is no baseline against which such a figure could be recomputed. That still holds, and this page is not an exception to it. The composite at the top is scaled against this project's own history and says nothing about anyone else's; the raw counts of our own activity are recomputable by anyone holding the same command, the same window, and the same commit. None of it asserts that the work was faster, cheaper, or better than any alternative, and no figure here should be read as making that assertion.
This is one repository, over one window, at one instant. It is not a fleet roll-up, not a benchmark, and not a comparison against anybody. Counts move while they are being taken, because the corpus changes during the measurement, so every number on this page is stamped with the instant and the source commit it was taken at, and means nothing without them.
One reference, and it is us. There are no external deployments yet. Every operational number here comes from our own fleet, the toolkit's own development, and it is labelled as such. We run our own company on it: one human, an agent fleet. It is the only reference that exists today, and you can read all of it. The same pipeline seen as a live board is on the board.
The flow metrics underneath the score
These are the per-brief flow measurements the score is built out of, plus the ones published beside it as context. When the composite moves, this is the table that says which way and why. Each is in the same three states as everything else on this page, measured, partial, or could-not-check, and a metric that could not be computed keeps its row and prints those words rather than showing a zero.
| Metric | Value | State | What it counts |
|---|---|---|---|
| Throughput throughput | 976 effort points · 344 briefs · 264 of unknown size | measured | briefs delivered in the window, and their weighted effort points |
| Lead time by size leadtime | S: 26.7 d median, 34.8 d p85, n = 103 · M: 17.7 d median, 36.8 d p85, n = 203 · L: 14.7 d median, 29.1 d p85, n = 30 | measured | median and 85th-percentile days from authored to done, split by brief size |
| Flow efficiency flow_efficiency | 0.9% of elapsed time worked · 12 working spans against 1232 waits | measured | the share of elapsed time a brief was actually being worked rather than waiting |
| First-pass yield first_pass_yield | 86.7% · 13 of 15 briefs · 45 excluded as unlinked | measured | the share of briefs that reached done without a rework round |
| Review rework review_rework | 0.40 review rounds per brief · n = 40 | measured | mean review rounds per brief, and how the rounds are distributed |
| Decision latency decision_latency | p50 74.6 hours · p90 429.7 hours · 22 open now · oldest waiting 523.2 hours | measured | how long an open decision waits on a human, and how many are open now |
| Net flow and stalls stall | could-not-check | could-not-check | the per-stream net-flow / stall feed was not probed for this snapshot |
The five delivery metrics
The commodity delivery family is published alongside, from its own feed. It is not folded into the composite. Each metric is in exactly one of three states. measured means the emitter computed it from a complete input. partial means it computed it from an input it itself declares incomplete, and the caveat travels with the number. could-not-check means it could not compute it at all. A could-not-check metric is never shown as zero and never dropped from the table: a zero for a probe that did not run is a false claim, and a dropped row hides that this page's coverage is partial.
| Metric | Value | State | What the emitter reports about it |
|---|---|---|---|
| Deployment frequency deployment_frequency | — | could-not-check | the feed did not carry this metric at all |
| Change lead time change_lead_time | — | could-not-check | the feed did not carry this metric at all |
| Failed-deploy recovery time failed_deploy_recovery_time | — | could-not-check | the feed did not carry this metric at all |
| Change failure rate change_failure_rate | — | could-not-check | the feed did not carry this metric at all |
| Rework rate rework_rate | — | could-not-check | the feed did not carry this metric at all |
Coverage of this table right now: 0 of 5 measured, 0 partial, 5 could-not-check — the commodity delivery (DORA) and velocity metrics were split out of this generator to a dedicated open-source metrics platform (Apache DevLake) and are served from there. They are not yet wired back into this snapshot, so the DORA series is UNREAD here — never a zero that reads as measured.. What cannot be computed is not a broken measurement. It is a measurement with no independent input yet. Publishing those as blanks would imply the system had nothing to report; publishing them as zero would imply it had nothing to fix. The last column is the emitter's own wording, kept verbatim rather than paraphrased, because a paraphrase would put a second author between you and what was reported. It names the desks (the standing roles that run the pipeline) that would supply the missing inputs; those roles are described in the methodology overview.
Trend over time
A trend view rolls the recorded lifecycle transitions up into a time series. Its state for this snapshot: could-not-check
the trend feed was not probed for this snapshot
Where each number comes from, and where it stops
Some inputs are list calls with a hard result cap. A count taken from a capped list call becomes a silent undercount the moment the cap binds: the call still reports success, and nothing downstream can tell the difference. The emitter's output does not state its own caps, so they are recorded explicitly when the snapshot is refreshed and published here alongside the numbers they bound.
| Probe | Bounds | Result cap | Behaviour at the cap |
|---|---|---|---|
| This snapshot states no probe caps, because the capped list calls behind the delivery table were not run for it. | |||
Raw activity
Commits per day, merges per day, and pull requests per day used to be the headline of this page. They are not any more, and they are not gone either. They are here, one fold down, because deleting a number you have stopped leading with is a different and worse act than demoting it.
Show raw activity: commits, merges, and pull requests per day
| Counter | Per day | State | Window |
|---|---|---|---|
| Commits | 140.07 | measured | trailing 28d (2026-08-09 to 2026-09-06) |
| Merges | 36.11 | measured | trailing 28d (2026-08-09 to 2026-09-06) |
| Pull requests | 33.86 | measured | trailing 28d (2026-08-09 to 2026-09-06) |
raw activity is shown only behind a fold; the headline is the AssayScore. Commits, merges and PRs count motion, not delivered value.
A repository can produce a great many commits per day while shipping less than it did last month, and any of these three numbers can be raised on purpose by anyone who decides to raise it, which is why none of them is the metric of record here. They stay published because they are cheap to recompute and easy to check, and because a page that quietly dropped its old headline the day it became inconvenient would be making a claim of its own.
What this page does not cover
- One repository only. The toolkit's own development repository. No other repository contributes a single number here, and nothing on this page should be read as a fleet-wide or organisation-wide figure.
- One window only, the dates stamped at the top. Nothing before the window start is counted; a longer or shorter window would produce different numbers, and no window here was chosen to flatter a result.
- No cross-organisation comparison. The composite is scaled against this project's own trailing history, so it is not comparable with a score computed against a different project's history. Publishing it as though it were would be the exact overclaim this page exists to avoid.
- Deployments are not production deployments. The delivery throughput figure counts commits and merges in the repository, which is what the emitter can see. It is not a count of releases reaching users, and should not be read as one.
- Change failure is a partial signal. It is computed from bug-labelled issues against merged changes. It does not include failures that were never filed as a bug-labelled issue, so it is a floor rather than a rate.
- Not everything in the feed is drawn here. The published snapshot carries additional diagnostic blocks, the constraint stage, the intake front-door backlog, and the human-decision queue, that this page does not yet render. They are in the feed and readable there, and the board page draws them; saying so is cheaper than pretending this page is the whole of it.
- No per-person anything. These are system-level aggregates. Nothing here is attributed to an individual, and the toolkit's own framing forbids using any of it as an individual scorecard.
The source this page was rendered from
Every number above is derived from one declared source, and nothing on this page is typed by hand. That source is published as metrics.assay.guide/metrics.json: the emitter's raw output, verbatim, plus the snapshot's commit, window, and probe caps. This page carries no numbers of its own. It is a static shell that fetches that feed live and renders it in your browser, so a fresh snapshot appears without redeploying the site. If the live feed cannot be reached it falls back to the copy committed next to the page.
The snapshot is regenerated on a schedule by the toolkit's own emitter and uploaded to the feed; the site itself is not rebuilt to move the numbers. A check re-renders the page shell and requires it to be byte-identical to what is published, which proves no number was hand-typed into the markup, and separately requires the published snapshot to be well-formed. This is the same stance the toolkit takes everywhere else: one writer, one source, and a script that can tell when the two have drifted.
The emitter's own framing, verbatim: the commodity delivery (DORA) and velocity metrics were split out of this generator to a dedicated open-source metrics platform (Apache DevLake) and are served from there. They are not yet wired back into this snapshot, so the DORA series is UNREAD here — never a zero that reads as measured.