Concept register · Theme 13 of 14 1 concepts · 14 talks
assay · concepts · autonomous-discovery
Autonomous discovery
Agent systems pointed at the whole research loop — ingest, hypothesize, implement, run real experiments, falsify most of it, write up what survived — with the human posing the question and accepting the result.
assay: shaped like the dreaming pass · unbuilt
§1What it is
Three infrastructural legs
The systems that work stand on auto-reset (a fresh environment per attempt), auto-improve (the loop drives its own upgrades) and auto-evaluate (outcomes judged by the environment, not by model confidence). Feedback must come from the environment, and selection matters as much as generation — the loop needs a mechanism for choosing among candidate ideas, not only for generating them.
Two dissents, recorded
The reported successes are mostly bounded, verifiable problems — exactly the class a register-and-gates methodology fits — and full autonomy over open-ended questions remains undemonstrated. The dissents are held, not dismissed.
§2The concepts in this theme
Each concept has its own page in the concept register — with sightings from every event we review, and where Assay stands on each.
- AI co-scientists and autonomous discovery loops established 11
§3How Assay implements this
The legs map onto what ships
Auto-reset is worktree discipline — every worker gets a fresh worktree off the mainline. Auto-improve is the worker desks. Auto-evaluate is the verify rows executed after merge by someone other than the implementer. The end-to-end research pipeline — out-of-band idea generation, agents arguing the candidates, harness-gated execution, human accept-or-reject at the end — is the shape the dreaming pass was designed to have. None of that runs today.
The strongest outside validation in the scan
The open-source “lab” results — branch per idea, sandboxed writes, findings written down — are Assay’s core bet arrived at independently, sandbox rationale included, with loops carrying all three legs saturating their task. Three transfers are recorded: an ELO-style tournament as a selection step among competing brief plans (none exists today), iterating with the environment as the argument for verify rows against real CI, and revisiting brief granularity on the task-horizon cadence rather than treating it as fixed.
§4Talks that cover this theme
#004A Lab Notebook for AgentsChuan Li, Lambda
#007From Models to Agents to DiscoverySaurabh Tiwary, Google
#017Opportunities and Challenges for Long Horizon AgentsJerry Tworek, OpenAI
#023Robotics: EndgameJim Fan, NVIDIA
#034The Eureka Machine: Recursive Superintelligence for ScienceRichard Socher, Recursive
#036Combining Experiments, LLMs, and Theory to Discover Quantum MaterialsEkin Dogus Cubuk, Periodic Labs
#079Unlocking Scientific Abundance by Learning from Superhuman AIEric Ho, Goodfire
#082Workshop: Open Source Agent InvestigationsLambda / Berkeley RDI
8 of 14 talks shown — the ones that reach the most concepts in this theme. Every sighting, per talk, is on the concept pages above.