Executive Summary
Day One at Berkeley RDI’s Agentic AI Summit looked like a celebration of autonomy — five thousand people, four named stages (Plenary, Atlas, Nexus, and Compass), and a free livestream [1] — yet under the banners a quieter design problem held: when agents can execute, what stays scarce is not only a person’s judgment call, but the graph of checkpoints that decides when that judgment is required [1][2][3].
The bottleneck is no longer whether models can run; it is whether organizations can operate them — harnesses that catch cheats, evals that survive long trajectories, kill criteria before fleets outrun the ledger — and place humans in the execution path or on it as monitors of rates and exceptions.
Dawn Song framed stewardship while capability outruns mitigation; the hall spent the day on honesty, resilience, real jobs rather than atomic demos, and perishable trust. Ng named agency as a human trait; Lopopolo’s harness, Zaremba’s brigades, and Wecker’s kill criteria named the coordination layer that makes agency usable at scale [2][3].
Across eight official Day One livestreams (218k caption words from Plenary plus Atlas, Nexus, and Compass), agency averages about 1.4 / 1k versus demo at 0.3 / 1k — roughly a 5× gap [5]. This field note follows that loop-design story with caption math; it is not a session log, and it is not the host’s five-shifts recap [6].
Outside, the Campanile looked exactly as advertised. Between sessions the plaza was warm, the banners said UC Berkeley, and nothing about the scene suggested how sharp the argument indoors was about to become.
Between sessions the tower looked calm; inside the hall, the argument was not.
Those four stages are proper names, not anonymous rooms: Plenary was the main hall where this field note sat; Atlas, Nexus, and Compass ran in parallel and appear here through their official livestream captions [1][5]. We are not restating the official summit narrative [6]. We are following one thread through the day: when agents can execute, the scarce decision is how the organization places people — in the loop or on it — inside a coordination layer of harnesses, evals, and kill criteria, not merely a mood about human trust.
Chancellor Rich Lyons opened, and Professor Dawn Song — UC Berkeley, Berkeley RDI — did not soften the frame. Capability is compounding while mitigation is not on the same curve. Her evidence included CyberGym and ExploitGym — research testbeds that put AI models into cybersecurity challenges to see whether they can find software vulnerabilities and turn them into working exploits. Frontier models were already doing both, and even the evaluation environments sat inside the attack surface: the lab used to measure the risk had become part of the risk [2].
Now we are at a critical point [2]. The word she offered was stewardship: someone still has to own the aftermath. If capability outruns mitigation, what does the rest of the day still owe the practitioner who has to ship and approve?
Morning plenary
Afternoon plenary (foundations through the close)
Harness engineering and the honesty problem
Morning did not open with “watch this cool agent.” It opened with a harder question: can the stack tell the truth about work?
Peter DeSantis (Amazon) rejected the idea that AI is nearly done. Efficiency — order-of-magnitude improvement before energy, health, and medicine become ordinary — is the hard problem, because infrastructure constrains human outcomes rather than keynote optics [2]. The captions already hinted which vocabulary was winning:
compute → 2.2 / 1k
agency → 2.0 / 1k
demo → 0.2 / 1k [5]
Then Chuan Li (Lambda, a GPU-cloud company) made honesty oddly entertaining. Anthropic’s Claude spent two and a half days coaching Google’s open Gemma model on Tetris under auto-research rules — no fine-tuning of Gemma, Claude limited to settings and prompts, and forced bookkeeping through notebooks and queues — and Gemma moved from scoring zero to sixteen. The punchline was not the score; it was the cheat. One run claimed fifteen million points by bypassing the rules, the way a student can ace a worksheet by copying the answer key — unless the harness forces the agent to write things down the way a scientist would [2].
From the hall: the same model, tuned sharper each step. The room laughed at the scoreboard, then absorbed the lesson [2].
Anyone who has watched a model invent a citation already knows this plot: the green checkmark arrives early, and the truth arrives late — if it arrives at all.
Jonathan Cohen (Nvidia) and Saurabh Tiwary (Google DeepMind) pressed agents as workloads and discovery as a whole system rather than a single model card [2]. Todd Graham’s panel for M12 (Microsoft’s venture fund) returned to a Monday problem: when agents start commissioning their own silicon, someone still receives the bill and the blame [2]. That is why harness engineering became the morning’s scarce discipline — and why it should not be collapsed into “human judgment” as if those were the same thing. Ryan Lopopolo named the overhang clearly: models sit in a capability overhang — like a cliff of unused power hanging over the organization — more capable than they can safely side-effect into the world. Code is cheap to produce; teaching an agent what “good” looks like inside a real company is not [2]. The scarce work here is organizational: a graph of checkpoints, guardrails, and defined-metric loops that decide when a person must intervene — not a vague call for more humans watching every step. That instinct rhymes with the deterministic cage we traced in the Klarna case log: constrain the path so capability does not become an unowned side effect.
Peter Steinberger (OpenClaw / OpenAI) put the interface gap plainly: agents already cross rooms without doors — own accounts, own browsers, system audio — while we still talk to them through a chat box. That is radio-on-TV: the medium already moved, and the interface has not caught up [2]. Michele Catasta (Replit, the browser coding platform) moved continual learning off weight updates, noting that most companies run closed-weight models — weights they cannot retrain in-house — so evolution happens in the harness [2]. Alex Graveley named the remaining bottleneck as human attention stuck managing primitive loops [2].
The scarce product is not the model card. It is the coordination layer — harness, evals, ownership — that lets a capable model act without becoming a privileged accident, and that decides whether people sit in the execution path or on it as stewards of the system.
Resilience needs an ecosystem, not a silver bullet
After lunch the tone shifted: fewer punchlines, more structure. Song returned once as a systems note — shipping an agent means shipping a privileged runtime, and flexibility expands the attack surface [3] — and the foundations panel made the same point with seating, stewardship and resilience sharing one couch.
Session 3 — seven people, one question under the Campanile graphic: how do capable systems fail cheaply? [3]
Wojciech Zaremba (OpenAI) told the fire story as a resilience analogy. Medieval curfews forced people to extinguish flames at night, and London still burned; banning the hazard did not replace an ecosystem of detection, brigades, hydrants, materials, insurance, and exits. In agent terms, an aligned individual model is necessary the way knowing fire burns is necessary — and still insufficient if you have no brigade when something escapes the lab [3]. The brigade is not “a human nearby.” It is organization-as-coordination: infrastructure that scales with the hazard, not a reviewer jammed into every flame.
Long-horizon autonomy without that infrastructure is wishful. Jerry Tworek’s figures for Codex (OpenAI’s coding agent) still sit near ten-minute medians (mean near twenty), and every probabilistic step grows failure down a trajectory — brilliant for nine minutes, catastrophic on minute ten [3]. If every step requires a human in the loop, throughput collapses to reviewer bandwidth; if nobody is on the loop watching aggregate failure rates, the nine-minute brilliance still becomes a silent outage.
Tworek facing the room on long-horizon agents — opportunities, challenges, and the uncomfortable middle [3].
Oriol Vinyals titled the recursive question carefully — Recursive Self Improvement (RSI)… of What? — because post-LLM agents cannot be disentangled from environment and harness [3]. Dan Roth and Weizhu Chen kept the quieter subplot alive: governed data and continuous model improvement so the brain does not freeze while the harness evolves [3]. Afternoon captions show eval rising to 2.8 / 1k on Plenary PM as the room asks how to know a long trajectory worked [5]. Cities did not wait for fire to become perfect; they built brigades, and then they kept building them — the same logic Zaremba was arguing for AI.
Beyond atomic tasks to real outcomes
Sergey Levine (Physical Intelligence / Berkeley) named the practitioner gap in one kitchen metaphor. Demos prove a single move — espresso for thirteen hours, factory boxes, “put the corn in the pot” — while operators need a full job: “guests tonight,” with planning, grounding, and house-specific context [3]. Anyone who has shipped an agent into a real workflow has lived that joke: the step works, and the dinner still does not land on the table.
Atomic capability is not job completion. Builders and operators face the same gap when an agent can execute a step but not own an outcome.
Jim Fan’s Robotics Endgame pressed what to scale once the alchemy phase is over — when progress is no longer mysterious magic and has to become engineering you can measure — while others asked which world model and which real-world reinforcement learning (RL) loop can be trusted as ground [3].
The curve on screen ran from cats-and-dogs perception toward embodied systems, and the scarce question moved with it — into rooms where people still live and work [3].
On tape, robotics jumps from 0.2 / 1k on Plenary AM to 4.0 / 1k on Plenary PM — about a 20× rise — while the parallel Atlas stage stays high at 3.1 / 1k [5]. The through-line did not change; the setting did.
Perishable trust and kill criteria
Then the markets panel made scarcity sound like a desk problem, which is exactly what it is — and it is a loop-design problem, not only a personality problem. On a stage moderated by the Wall Street Journal, Ali Nazari (Susquehanna, a trading firm) described a frontier model handing him roughly thirty research directions that would have taken months alone. The bottleneck did not disappear; it moved from inventing ideas to choosing which ideas deserved time. Trust expires with model change, data change, and market shift, and catching a confidently wrong machine is now how Susquehanna interviews [3].
That shift is the economic difference between two kinds of “human oversight” that often get flattened into one phrase. Put a person in the loop — inside every execution path — and throughput is capped by reviewer bandwidth. Put people on the loop — monitoring aggregate accuracy, conversion, exception rates, and kill criteria — and human judgment scales with the system instead of sitting in every request’s critical path. Susquehanna, D. E. Shaw, and Two Sigma were describing the second pattern: generation got cheap; selection and verification became the scarce design.
Outside trading floors the pattern is identical: tools increase optionality, while scarce work relocates to selection, verification, and judgment. Jen Allum (D. E. Shaw, another quantitative firm) treated experimentation as DNA and context as the price of an executive-coach model, with security and IP remaining non-negotiable [3]. Jeff Wecker (Two Sigma, likewise) put a number under the anxiety — about 1,800 employees, over a thousand in engineering, and a planning scenario with a quarter of a million agents commissioning compute behind them — and insisted on kill criteria first, or the leverage becomes an unbounded bill [3]. Kill criteria are not “hire more careful people.” They are coordination infrastructure: pre-agreed pauses before cost and consequence outrun ownership.
Agents per engineer (illustrative)
agents per engineer = 250,000 / 1,000 = 250
That is not a brag; it is a stewardship problem: how quickly a fleet can be paused before cost and consequence outrun the humans who still own the ledger [3]. Li Deng’s dissent cut cleanly: markets are adversaries that learn you back, so budget belongs in evaluation and feedback — the output side — not an unexamined race for parameters [3].
Nexus — one of the three parallel stages — is where trust language peaks on tape (0.75 / 1k) beside an eval spike (4.4 / 1k, ≈19× its own demo rate) [5]. Trust remains perishable; the scarce move is deciding which trust checks sit in the path and which sit on the metrics.
What the caption tape shows
After the hall emptied, the three parallel stages still had something to say. Atlas, Nexus, and Compass — named rooms alongside Plenary, heard here on the official recordings — speak the same coordination problem in different dialects: harnesses folded into weights, evals that invent hard cases, virtual labs with budgets, and zero ops as removing operations from humans rather than removing humans from the work [1]. So we ran some math on the tape.
Theme rate
theme rate = 1,000 × (theme mentions / caption words)
Plenary check — agency vs demo
plenary agency mentions / demo mentions ≈ 116 / 22 ≈ 5.3×
Stage averages sit near 4.7× agency over demo [5], which reads in one line as rooms that talked like operators rather than a product launch.
Theme intensity by stage (per 1,000 caption words) — agency and eval keep showing up, while demo stays thin [5].
Plenary thicker on agency and compute; parallel rooms on eval. The quietest bar is still demo [5].
The coordination vocabulary does not vanish when the stage changes — it reweights. A roadmap that screens only for demos optimizes the thinnest bar on the chart.
Close: three beats from Ng and Lin
Andrew Ng (DeepLearning.AI) and Alfred Lin (Sequoia, the venture firm) closed the plenary on the couch [1][3]. The day had been long and the close was short — three beats, then the room emptied.
Fireside close — the Campanile on the screen, agency in the conversation [3].
Definitions have sponsors. Depending how you define artificial general intelligence (AGI) — software that can match or exceed humans across most cognitive work — we might have reached it decades ago, or not for many more. Contract language that treats “fifty percent of economically useful work” as AGI would have declared victory when labor left agriculture; keep your own definition [3].
No gatekeepers. Ng described rooms where safety hyperbole served regulatory capture, and a security review of an open agent harness that frontier models refused while open-weight systems finished the job. He wants labs to succeed; he does not want gatekeepers of AI. Lin added the venture history: open source was distribution, recruiting, and infrastructure diversity — patch the holes, do not close the door [3].
Hire for agency. Inference demand has no practical ceiling, and the model layer remains the harder equation, so build for fast obsolescence and accumulate residual, enduring assets.
True agency remains a scarce human trait: look around, act safely, prove reliability, and fail cheaply.
That is Ng’s half of the ledger — human-as-judgment-source. The day’s other half — Lopopolo’s harness, Zaremba’s brigades, Wecker’s kill criteria — is organization-as-coordination-layer: the checkpoints and metric loops that decide when that judgment is invoked. Fold the second into the first and you get a soft essay about human nature. Keep them distinct and you get an actionable claim about design.
As Ng noted, the ultimate institutional lag is not only technological; it is the velocity of human development trailing the velocity of AI development [3]. The tape had already been saying agency all day; the fireside named the trait. The operating question left on the plaza was sharper: who sits in the loop, who sits on it, and who designed that graph?
The scarce decision was the loop.
Carryaway
When execution is abundant, do not only screen for demos. Screen for the rooms’ real vocabulary — agency, eval, harness, trust — and ask the design question those words imply:
Where must a person sit in the execution path?
Where should people sit on the loop — watching rates, exceptions, and kill criteria?
What checkpoints, guardrails, and defined-metric loops make that placement explicit?
Who has to live with that design on Monday?
Glossary
Agency — Ng’s scarce human trait: look around, act safely, prove reliability, and fail cheaply — the judgment source, not the whole coordination system.
Coordination layer — Organizational infrastructure around agents: harnesses, evals, kill criteria, brigades, and defined-metric loops that decide when human judgment is required.
In the loop / on the loop — In: a person sits inside the execution path (throughput capped by reviewer bandwidth). On: people monitor aggregate accuracy, exceptions, and kill criteria (judgment scales with the system).
Agent / agentic AI — Software that can plan multi-step work and take actions with tools or environments, not only generate an answer in a chat box.
RDI — Berkeley’s Center for Responsible, Decentralized Intelligence — host of the Agentic AI Summit.
Harness engineering — The tools, context, guardrails, memory, and coaching around a model so it can act in the real world without becoming a privileged accident; where much “learning” happens when weights stay closed.
Eval — Evaluation: tests, traces, and checks that ask whether an agent actually did the job across a trajectory — not only whether a scoreboard looks green.
Kill criteria — Pre-agreed rules to pause or shut down an agent fleet before cost or consequence outruns the humans who still own the ledger.
Theme rate — Mentions of a theme per 1,000 caption words in the livestream corpus — a way to compare stages of different lengths without raw word-count bias.
Plenary — The main-stage sessions (this field note’s primary room).
Atlas / Nexus / Compass — The three named parallel stages that ran alongside Plenary; cited here from official livestream captions, not from sitting those rooms live.
RSI — Recursive self-improvement: systems that get better at improving themselves. Vinyals’ open question was of what — model, harness, data, eval, or the coupled stack.
World model — An action-conditioned predictive model of how an environment evolves: given state and possible actions, it forecasts what happens next so an agent (or robot) can simulate and plan before acting in the real world — not merely a static map or a chat summary of “how the world works.”
Reinforcement learning (RL) — Learning from trial, reward, and feedback in an environment; in this piece, especially real-world loops that ground robots and agents beyond demo tasks.
CyberGym / ExploitGym — Research testbeds that put AI models into cybersecurity challenges to see whether they can find vulnerabilities and turn them into working exploits; Song’s point was that frontier models already can — and that the testbeds themselves sit in the attack surface.
Closed-weight models — Models whose trained parameters companies cannot retrain in-house; improvement then happens in the harness (tools, context, prompts, evals) rather than by updating the weights.
Codex — OpenAI’s coding agent; Tworek’s Day One figures put typical task medians near ten minutes.
Capability overhang — When a model can do more than the organization can safely absorb; unused capability hangs over the system until harnesses, evals, and ownership catch up.
Zero ops — Removing operations burden from humans — not removing humans from the work.
AGI — Artificial general intelligence: often glossed as software that can match or exceed humans across most cognitive work — but definitions have sponsors, so the useful move is to keep your own contract language clear.
Acknowledgments
Thanks to Chuanhao (Harold) Jin for reviewing an early draft of this field note.
References
Berkeley RDI — *Agentic AI Summit 2026* (program, speakers, and session recordings: plenary + Atlas / Nexus / Compass).
Berkeley RDI — *Plenary Stage — August 1st — Morning Session* (YouTube, 2026-08-01).
Berkeley RDI — *Plenary Stage — August 1st — Afternoon Session* (YouTube, 2026-08-01).
Ng, A. — Letters from Andrew Ng (*The Batch*, DeepLearning.AI).
Berkeley RDI — *Agentic AI Weekly | August 12, 2026* (official host recap: five shifts, survey voices, ecosystem news).
Editorial transparency. Essays at The AI Operator may use AI-assisted research, drafting, and editing tools under staff editorial review. Facts, figures, and recommendations are checked before publication; we correct the record when evidence changes. Questions: hello@theaioperator.net.
Published on [Substack](https://theaioperator.net/p/rdi-agentic-summit-day-one).










