Argues that agent harnessThe scaffolding around a model that turns it into a working agent — the loop, the tools it can call, and the rules for when to stop.Full definition → design, not just model capability, limits coding agents, and explains RLM techniques like context windowThe maximum amount of text a model can consider at once — its working memory for the current conversation or task.Full definition → offloading and persistent subagents you can apply when orchestrating agents.
Key takeaways · AI-distilled
Zhang argues Claude Code, Codex, Pi and many other modern agent harnesses are structurally very similar.
The episode cites OpenAI's 10,000-agent experiment, about 130B output tokens and roughly $40M-equivalent spent on one problem, and notes much of a swarm's work may be wasted search.
Zhang says AI-generated GPU kernels still leave substantial room for human expertise, and one expert insight can potentially replace large amounts of brute-force token search.
Speculative programmatic tool useA model's ability to call external functions — run code, search the web, edit files — instead of only generating text.Full definition →, which overlaps tool execution with generation, is one of the techniques Zhang discusses for faster agents.
Terms in this piece · Glossary
multi-agent — Using several AI agents on one problem — splitting work in parallel, checking each other, or filling different roles like planner and reviewer.
agent harness — The scaffolding around a model that turns it into a working agent — the loop, the tools it can call, and the rules for when to stop.
context window — The maximum amount of text a model can consider at once — its working memory for the current conversation or task.
tool use — A model's ability to call external functions — run code, search the web, edit files — instead of only generating text.