Advanced Workshop: Mastering AI Observability — Doug Guthrie, Braintrust
Source
youtube.com
Author
AI Engineer
Date
Why it matters
Walks through closing the loop from production traces to evaluation datasets to code changes, so agent regressions are caught against real failure cases.
Key takeaways · AI-distilled
The workshop's demo is a Python support AI agentAn AI system that doesn't just answer once but works toward a goal in a loop — taking actions, reading the results, and deciding what to do next.Full definition → built on the OpenAI Agents SDK. A batch of 968 sample support traces surfaces order lookup failures and broken escalation calls as the cases worth investigating.
Online automations grade incoming traces while Braintrust's Topics feature clusters patterns in tasks, sentiment and workflow issues. Guthrie adds a custom facet to catch failures that the default issue classification misses.
Braintrust's embedded Loop assistant queries trace data with SQL to examine representative failures, and selected examples go back into evalA repeatable test for AI quality — a set of tasks plus scoring — used the way software teams use test suites, because model output is too variable to judge by eyeballing.Full definition → datasets so later changes are checked against real failure cases.
A closing demo exposes an agent evaluation to a playground, so collaborators can adjust prompts and models while execution stays on the evaluation server.
Terms in this piece · Glossary
AI agent — An AI system that doesn't just answer once but works toward a goal in a loop — taking actions, reading the results, and deciding what to do next.
eval — A repeatable test for AI quality — a set of tasks plus scoring — used the way software teams use test suites, because model output is too variable to judge by eyeballing.