Vibeleaderboard
← All Intel
Intel / video

From Vibes to Production: Evaluating and Shipping AI Agents That Work 101 — Laurie Voss, Arize AI

Source
youtube.com
Author
AI Engineer
Date
Why it matters

Shows an judge failing every report because it lacked the agent's research , then fixing it by supplying sources. Gives a concrete workflow for tracing, rubric design and validating judges against human labels.

Key takeaways · AI-distilled
  • Supplying the 's collected research sources to a faithfulness evaluator turned a judge that rejected all 13 financial reports into a usable split of six faithful and seven unfaithful reports.
  • Reading OpenInference traces of the Claude Agent SDK agent exposed failed file writes, missing report content and excessive web searches.
  • Laurie Voss layers : a cheap deterministic ticker check first, then evaluators that each check one dimension with explicit criteria, tagged inputs, examples and binary labels, so a failure names what to fix.
  • Judges are compared with human annotations using precision and recall and checked for biases. Failing traces then become datasets, and controlled experiments compare the changed agent on the same cases.
Terms in this piece · Glossary
  • AI agent — An AI system that doesn't just answer once but works toward a goal in a loop — taking actions, reading the results, and deciding what to do next.
  • context window — The maximum amount of text a model can consider at once — its working memory for the current conversation or task.
  • eval — A repeatable test for AI quality — a set of tasks plus scoring — used the way software teams use test suites, because model output is too variable to judge by eyeballing.
  • LLM — A large language model — the neural network behind tools like Claude and ChatGPT, trained on huge amounts of text to predict what comes next.
Read the source www.youtube.com
More from AI Engineer
Recommended reads
Comments

Checking sign-in…

Loading comments…