From Vibes to Production: Evaluating and Shipping AI Agents That Work 101 — Laurie Voss, Arize AI
Source
youtube.com
Author
AI Engineer
Date
Why it matters
Shows an LLMA large language model — the neural network behind tools like Claude and ChatGPT, trained on huge amounts of text to predict what comes next.Full definition → judge failing every report because it lacked the agent's research context windowThe maximum amount of text a model can consider at once — its working memory for the current conversation or task.Full definition →, then fixing it by supplying sources. Gives a concrete workflow for tracing, rubric design and validating judges against human labels.
Key takeaways · AI-distilled
Supplying the AI agentAn AI system that doesn't just answer once but works toward a goal in a loop — taking actions, reading the results, and deciding what to do next.Full definition →'s collected research sources to a faithfulness evaluator turned a judge that rejected all 13 financial reports into a usable split of six faithful and seven unfaithful reports.
Reading OpenInference traces of the Claude Agent SDK agent exposed failed file writes, missing report content and excessive web searches.
Laurie Voss layers evalA repeatable test for AI quality — a set of tasks plus scoring — used the way software teams use test suites, because model output is too variable to judge by eyeballing.Full definition →: a cheap deterministic ticker check first, then evaluators that each check one dimension with explicit criteria, tagged inputs, examples and binary labels, so a failure names what to fix.
Judges are compared with human annotations using precision and recall and checked for biases. Failing traces then become datasets, and controlled experiments compare the changed agent on the same cases.
Terms in this piece · Glossary
AI agent — An AI system that doesn't just answer once but works toward a goal in a loop — taking actions, reading the results, and deciding what to do next.
context window — The maximum amount of text a model can consider at once — its working memory for the current conversation or task.
eval — A repeatable test for AI quality — a set of tasks plus scoring — used the way software teams use test suites, because model output is too variable to judge by eyeballing.
LLM — A large language model — the neural network behind tools like Claude and ChatGPT, trained on huge amounts of text to predict what comes next.