Brains vs Hands: How to Run AI Agents Safely in Production — Viren Baraiya
Source
youtube.com
Author
AI Engineer
Date
Why it matters
Let the model plan but not improvise execution. A durable harness with approval gates and idempotent, recorded steps makes long-running agents safe to operate, shown by compiling an SRE agent's plan into a workflow.
Key takeaways · AI-distilled
Viren Baraiya, Orkes CTO and original creator of Netflix Conductor, splits agents into brain and hands: the LLMA large language model — the neural network behind tools like Claude and ChatGPT, trained on huge amounts of text to predict what comes next.Full definition → decides what should happen next, a deterministic durable agent harnessThe scaffolding around a model that turns it into a working agent — the loop, the tools it can call, and the rules for when to stop.Full definition → performs the actual work.
He argues production agents run on schedules, react to events, coordinate other agents and can run for months, so the harness around them becomes the application, much like a set of microservices.
Baraiya describes AI agentAn AI system that doesn't just answer once but works toward a goal in a loop — taking actions, reading the results, and deciding what to do next.Full definition → harnesses as late-bound sagas: workflows the agent assembles at runtime, then executed with human approval gates, idempotent steps and recorded side effects.
In the demo, an SRE agent's plan is compiled into a Conductor workflow, executed, checked and re-planned across two loops.
Terms in this piece · Glossary
LLM — A large language model — the neural network behind tools like Claude and ChatGPT, trained on huge amounts of text to predict what comes next.
agent harness — The scaffolding around a model that turns it into a working agent — the loop, the tools it can call, and the rules for when to stop.
AI agent — An AI system that doesn't just answer once but works toward a goal in a loop — taking actions, reading the results, and deciding what to do next.