Transcript
As AI systems evolve from chat interfaces into autonomous agents capable of reasoning, planning, and tool usage, traditional evaluation approaches are breaking down. Offline benchmarks and static datasets fail to capture the complexity, non-determinism, and operational risks of real-world AI systems operating in production environments. In this talk, I’ll share practical lessons and architectural patterns for building evaluation systems for agentic AI workflows at scale. We’ll explore how modern AI platforms are shifting from one-time benchmark testing toward continuous evaluation pipelines integrated directly into production infrastructure. Topics include: - Why offline evals fail for autonomous AI systems - Evaluating tool use, planning, reasoning, and multi-step workflows - Online vs offline eval architectures - Human-in-the-loop evaluation systems - Detecting drift, hallucinations, and unsafe behaviors - Building feedback loops for continuous improvement - Observability and telemetry for agentic workflows - Reliability metrics beyond model accuracy Attendees will leave with practical frameworks for designing scalable evaluation systems capable of measuring real-world AI behavior, reliability, and operational impact. Speakers: - Nishant Gupta (Meta Superintelligence Labs): Nishant Gupta is a Software Engineering Tech Lead at Meta Superintelligence Labs focused on building the training and inference AI Infrastructure. LinkedIn: https://www.linkedin.com/in/nishantgupta-ai/ GitHub: https://github.com/nishantgpt-lab