From Vibes to Production: Evaluating and Shipping AI Agents That Work 201 — Laurie Voss, Arize AI
Source
youtube.com
Author
AI Engineer
Date
Why it matters
Shows an improvement loop where traces reveal a missing price filter, a coding agent fixes it, and regression evals protect existing behavior. Useful for teams moving agent debugging beyond manual trace reading.
Key takeaways · AI-distilled
Skills installed in the workshop repo let a coding AI agentAn AI system that doesn't just answer once but works toward a goal in a loop — taking actions, reading the results, and deciding what to do next.Full definition → both add Arize AX tracing to the Wonder Toys app and query the resulting trace data.
The trace diagnosis looks past exceptions and HTTP errors to search quality problems, such as empty results and overly narrow filters.
Laurie Voss argues reading traces one by one fails at millions of traces, and evalA repeatable test for AI quality — a set of tasks plus scoring — used the way software teams use test suites, because model output is too variable to judge by eyeballing.Full definition → still leave people interpreting every failure. Arize's Signal adds a layer that groups recurring problems and suggests fixes.
In the demo, Signal findings can become GitHub issues, evaluation datasets, new evaluators or proposed pull requests, and Signal itself is traced and has its own evaluation suite.
Terms in this piece · Glossary
AI agent — An AI system that doesn't just answer once but works toward a goal in a loop — taking actions, reading the results, and deciding what to do next.
eval — A repeatable test for AI quality — a set of tasks plus scoring — used the way software teams use test suites, because model output is too variable to judge by eyeballing.