
If you're moving agents from demos to production, this lays out concrete architectural patterns for continuous — evaluating multi-step and planning, detecting drift and unsafe behavior, and building feedback loops that static offline benchmarks can't capture.
“The question is no longer did the model generate the right answer? The question is did the system behave correctly?”
Nishant Gupta
“Because benchmarks measure model capability. Production measures system behavior.”
Nishant Gupta
“Reliability becomes the North Star metric. Accuracy becomes the only input.”
Nishant Gupta
“Production traffic is no longer just traffic. It becomes evaluation data.”
Nishant Gupta
“Evaluation is no longer just a phase, it's an operational capability.”
Nishant Gupta
videoWhy AI Agents Need Million-Token Context — Thomas Wolf & Olive Song, MiniMax
videoYour company brain will leak secrets: how we stopped it for big banks — Tanmai Gopal, PromptQL
videoTethered: Our Agents Are Us — Shu Fang, Two Sigma
videoAgents' next frontier: agent-to-agent and network effects — Jean-Denis Greze, TownChecking sign-in…
Loading comments…