
If you're shipping custom agents for teams, this lays out a practical loop for keeping them reliable as complexity grows — using a builder to run , diagnose failures across /tool/code layers, and open PRs, while escalating to humans only when it can't safely resolve a failure.
“given that AI agents is just one type of software, as you may guess, we are using AI to build AI”
Alfonso Graziano
“You can see the golden dataset as a test suite, but in a non-deterministic scenario.”
Alfonso Graziano
“updating the golden data sets or the scorers just to let the evals pass is not a good idea”
Alfonso Graziano
“the coding agent found new ways that humans didn't find um to improve the agent, and we got plus 10% on some of our internal benchmarks”
Alfonso Graziano
“if we don't know what's happening when we ship in production, we are basically blind”
Alfonso Graziano
videoWhy AI Agents Need Million-Token Context — Thomas Wolf & Olive Song, MiniMax
videoYour company brain will leak secrets: how we stopped it for big banks — Tanmai Gopal, PromptQL
videoTethered: Our Agents Are Us — Shu Fang, Two Sigma
videoAgents' next frontier: agent-to-agent and network effects — Jean-Denis Greze, TownChecking sign-in…
Loading comments…