Concept register · Concept 62 of 64 · Theme: autonomous discovery Reviewed 2026-09-01

assay  ·  concepts  ·  autonomous-discovery

AI co-scientists and autonomous discovery loops

Agent systems are being pointed at the whole research loop rather than at one step of it: ingest the literature, generate hypotheses, implement, run real experiments, falsify most of them, and write up what survived — with the human moved up to posing the question and accepting or rejecting the result.

established · assay: designed, unbuilt

11 independent sources · sighted at the Agentic AI Summit 2026 · last reviewed 2026-09-01


§1What it is

Three infrastructural legs

The loop runs without a human in each turn only when three things exist: auto-reset of the environment, auto-improve of the artifact, and auto-evaluate of the result. Where all three are present, tasks have been driven from 0% to near-total success autonomously. Where one is missing, a human is back in every turn and the loop is a demo.

Feedback must come from the environment

A breakthrough is by definition outside the training set, so thinking harder about the literature does not substitute for running the experiment. The history cited against “thinkism” is pointed: the Nobel that mattered for superconductivity was for liquefying helium — a capability build — with the phenomenon stumbled into; a headline superconductor came from trying tens of thousands of materials rather than from theory. Mathematics and theoretical computer science are a different regime.

Selection matters as much as generation

The systems that work falsify aggressively, often by rating candidate ideas against each other in a tournament before committing compute. Progress arrives in jumps, not smoothly — one autonomous run had 240 experiments between its last two improvements — which makes cheap selection the difference between a loop that pays and one that burns.

Two dissents, recorded

A frontier lab reports seeing no particularly successful auto-research rollout with today’s models: progress appears early and then stalls, because research is serial and long-horizon. And the field’s own framing is that autonomy on grand challenges is gated by trust rather than by capability.


§2Sightings

Agentic AI Summit 2026 · 14 sightings

Also: AlphaEvolve; AlphaFold and AlphaGenome; the FrontierScience benchmark; PutnamBench; METR time-horizon measurements; Goodfire Silico; Prima Mente Pleiades; and “the lab” (open source — workspace, dashboard, one git branch per idea).


§3Where Assay stands

Designed, unbuilt — and the closest external match in the scan

The end-to-end research pipeline is the shape Assay’s dreaming pass was designed to have: out-of-band idea generation, agents arguing the candidates, harness-gated execution, and a human accept or reject at the end. None of that runs today.

The three legs map cleanly onto what is shipped

Auto-reset is worktree discipline — every worker gets a fresh worktree off the mainline. Auto-improve is the worker desks. Auto-evaluate is the Verify rows executed after merge by someone other than the implementer, per the lifecycle. The 0%-to-99% result is evidence that loops with all three legs saturate their task, which argues for closing the remaining gap rather than adding a fourth mechanism. The open-source “lab” is the strongest outside validation of Assay’s core bet — branch per idea, sandboxed writes, findings written down — arrived at independently, sandbox rationale included.

Three specific transfers

The ELO idea tournament is a candidate mechanism for choosing among competing brief plans or proposed memory updates, where Assay currently has no selection step at all. The iterate-with-the-environment stance is the external argument for why Verify rows run against real CI rather than against model confidence. And the time-horizon cadence is a standing review trigger: if reliably-executable task length doubles roughly every six months, the granularity ceiling on a brief should be revisited on that cadence rather than treated as fixed. The dissent is worth holding too — the successes reported here are mostly bounded, verifiable problems, which is exactly the class a brief is.


§4Watch

  • Whether any auto-research result is independently reproduced — the striking numbers here are each single-team single-run reports.
  • Whether the dissent softens: a long-horizon research rollout that does not stall would be the signal that the loop generalizes past bounded tasks.
  • Whether idea-selection mechanisms (ELO tournaments, verifier-fed replanning) show up outside research settings, which is what would make them portable to plan authoring.