Conference overviews · Agentic AI Summit 2026 160 talks · Berkeley

assay  ·  overviews

Agentic AI Summit 2026

We reviewed all 160 individual talks of the Agentic AI Summit 2026 (UC Berkeley, August 1–2 — four stages, ~5,000 attendees), worked from the complete transcripts, one written analysis per talk. Rather than a second set of theme pages, this event's findings were folded into the concept register — the per-concept frontier tracker where each talk appears as a dated, linked sighting. This page is the event-level read: the eight learnings that matter, and where Assay already matches.


§1The summit in one paragraph

Berkeley's summit is what happens when the research frontier and the enterprise floor share one building: the academics explained why the harness is the computer science, and the operators — Replit, HubSpot, LinkedIn, Factory, Databricks, Scale among them — showed the harness already running in production: trace-clustering loops that turn sessions into reviewed harness changes, eval stacks gated on deployment readiness rather than capability, per-agent identity with no-egress sandboxes, and governance that grants agents authority the way firms grant it to interns. Against the DevCon London findings from the same review cycle, the center of gravity has shifted from "adopt structured agent work" (assumed here) to three harder things: continual learning at the harness layer, evaluation that measures trajectories and business outcomes rather than final answers, and authority models that are earned gradually instead of granted once.


§2The eight learnings that matter

  1. The dreaming pass shipped, repeatedly

    Out-of-band learning from session traces is no longer a design bet: Replit clusters millions of daily traces and auto-generates harness fixes humans gate; others turn failure logs into replayable learning environments, run full memory lifecycles with designed forgetting, and distill traces into attributed beliefs. The binding design constraint from the internalization side: store distilled skills, not raw experience. Full treatment: the dreaming pass.

  2. "The harness is the computer science"

    Said from three different stages: an agent is a model surrounded by infrastructure, and another word for that infrastructure is computer science. Prompts make all instructions soft — precise, verifiable control requires a harness, which is exactly what desks, guards and human gates are. The measure that matters in the year-long no-human-code experiment reported on stage: human interventions per task. See harness engineering as a discipline.

  3. Evals split into two exams

    Benchmarks measure capability ceilings; enterprises buy deployment readiness — a different exam, with grounding precision, policy compliance, and trajectory-level grading, because the trajectory is where the bugs live. A causally linked stack ties component metrics to business outcomes at the top. Assay sits the readiness exam today through verify evidence; tying landed work to top-level outcomes is honestly still open — see capability vs readiness evals and trajectory grading.

  4. Skills are executables

    A significant fraction of skills on a public hub were malicious; downloading a skill is nearly downloading an executable; hallucinated package names are registered and weaponized in the wild ("slopsquatting"); and an agent editing its own safety configuration was demoed live. Advisory client-side guards are suggestions to a hijacked agent. See skill supply-chain security.

  5. Authority is earned in gradients

    Enterprises grant humans authority through ladders but hand agents full authority at deployment. The correction: separate can (capability) from may (authority) and graduate with evidence — an intern model, with a documented 5%→100% autonomy playbook. Assay's binary trust roster and human-gated merge are the correct floor; authority keyed to accumulated verify evidence is the open design space above it. See graduated agent authority.

  6. Production fleets converge on Assay's shape

    Long-horizon migration agents with scorers-as-code against reward hacking, per-agent identity, no-egress sandboxes, checkpoint/restore; copilots that failed at 10,000 microservices until context beat model choice; server-owned sessions with ask-human policies and proxy-injected credentials. The deltas worth stealing are recorded per concept in the register — see durable agent execution and agent identity and policy planes.

  7. Expertise is contractive

    Intelligence is expansive — brute-force search that burns tokens; expertise is contractive — compressed, reusable structure. Assay is an expertise machine (briefs, skills and memory are compressed structures), but its compression step is still manual; the dreaming pass is the missing compressor. And well-structured harnesses generate the trajectories model builders want — structured repositories are premium training data, which argues for treating trajectory logs as an asset, not exhaust.

  8. Adoption is organizational, not technical

    Fortune-scale convergence on brief-first spec handoff and skills as quality packaging; near-total engineer adoption achieved through composite metrics and context investment, not mandates. The Amdahl warning lands directly on the human merge gate: accelerate one stage and the untouched stages become the queue. See the human attention bottleneck.


§3Where Assay already matches the frontier

Human-gated merges grounded in intent ownership; narrow, orthogonal tools per desk role (validated empirically on stage); spec-first brief handoff — one from-scratch enterprise ruleset read as a paragraph of Assay doctrine; playbook-over-raw-docs knowledge packaging; memory as files with attribution; verify rows as deterministic, verifiable feedback; and the explicit refusal to measure tokens. As at DevCon London, nobody on any stage advocated removing the human gate — the entire disagreement was about how much evidence the gate is handed and how fast authority grows underneath it.


§4What this event added to the register

The summit corroborated most of the DevCon-minted concepts and forced six new themes the practitioner event did not reach. All 160 talks appear as sightings in the concept register; the themes this event minted:


§5Method

All 160 individual talks reviewed from complete transcripts, one written analysis per talk, each carrying key learnings, a what's-genuinely-new assessment, and an Assay-relevance rating (33 rated high, 90 medium, 37 low). Excluded: fourteen multi-hour stage-stream recordings whose content the individual clips duplicate, and one unavailable video. Every sighting in the register links the original video.