Concept register · Theme 13 of 14 1 concepts · 14 talks

assay  ·  concepts  ·  autonomous-discovery

Autonomous discovery

Agent systems pointed at the whole research loop — ingest, hypothesize, implement, run real experiments, falsify most of it, write up what survived — with the human posing the question and accepting the result.

assay: shaped like the dreaming pass · unbuilt

AI co-scientists and autonomous discovery loops — 11 sources, established
1 concepts · 11 independent sources · 1 established

§1What it is

Three infrastructural legs

The systems that work stand on auto-reset (a fresh environment per attempt), auto-improve (the loop drives its own upgrades) and auto-evaluate (outcomes judged by the environment, not by model confidence). Feedback must come from the environment, and selection matters as much as generation — the loop needs a mechanism for choosing among candidate ideas, not only for generating them.

Two dissents, recorded

The reported successes are mostly bounded, verifiable problems — exactly the class a register-and-gates methodology fits — and full autonomy over open-ended questions remains undemonstrated. The dissents are held, not dismissed.


§2The concepts in this theme

Each concept has its own page in the concept register — with sightings from every event we review, and where Assay stands on each.


§3How Assay implements this

The legs map onto what ships

Auto-reset is worktree discipline — every worker gets a fresh worktree off the mainline. Auto-improve is the worker desks. Auto-evaluate is the verify rows executed after merge by someone other than the implementer. The end-to-end research pipeline — out-of-band idea generation, agents arguing the candidates, harness-gated execution, human accept-or-reject at the end — is the shape the dreaming pass was designed to have. None of that runs today.

The strongest outside validation in the scan

The open-source “lab” results — branch per idea, sandboxed writes, findings written down — are Assay’s core bet arrived at independently, sandbox rationale included, with loops carrying all three legs saturating their task. Three transfers are recorded: an ELO-style tournament as a selection step among competing brief plans (none exists today), iterating with the environment as the argument for verify rows against real CI, and revisiting brief granularity on the task-horizon cadence rather than treating it as fixed.


§4Talks that cover this theme