Production Evals For Agentic AI Systems - Nishant Gupta, Meta Superintelligence Labs
www.youtube.com- Category
- AI Agents
- Type
- ARTICLE
- Added
- Jul 24, 2026
About
As AI systems evolve from chat interfaces into autonomous agents capable of reasoning, planning, and tool usage, traditional evaluation approaches are breaking down. Offline benchmarks and static datasets fail to capture the complexity, non-determinism, and operational risks of real-world AI systems operating in production environments. In this talk, I’ll share practical lessons and architectural patterns for building evaluation systems for agentic AI workflows at scale. We’ll explore how modern
What it can do
Continuously evaluate agentic AI workflows in production
Live agent execution traces from production infrastructure → Ongoing reliability and behavior evaluation results
Evaluate tool use, planning, and multi-step reasoning
Agent action sequences and tool invocation logs → Assessment scores for tool use, planning, and reasoning quality
Detect drift, hallucinations, and unsafe behaviors
Agent outputs and runtime telemetry → Alerts flagging drift, hallucinations, or unsafe actions
Route evaluations through human-in-the-loop review
Agent outputs requiring human judgment → Human-validated evaluation labels and feedback
Build feedback loops for continuous improvement
Evaluation results and human feedback → Improvement signals fed back into the AI system
Capture observability and telemetry for agentic workflows
Running agent processes and execution events → Telemetry traces and observability dashboards
Measure reliability metrics beyond model accuracy
Production agent performance data → Reliability and operational impact metrics
Why it made the leaderboard
If you're moving agents from demos to production, this lays out concrete architectural patterns for continuous evaluation — evaluating multi-step tool use and planning, detecting drift and unsafe behavior, and building human-in-the-loop feedback loops that static offline benchmarks can't capture.
Media

Comments (0)
No comments yet
Indexed by a proprietary survey. Corrections welcome.