
If you're building or grading AI agents, this lays out why static and evals break as agents gain reasoning and tool-use, and makes the case for agent-as-a-judge systems that surface unanticipated failure modes and can even open PRs to fix them.
“Evals have gone from the new skill that every PM and every AI engineer has to learn to the thing that every serious AI team is betting on.”
Aparna Dhinakaran
“Every one of these was actually a massive jump in complexity, and we didn't just make the problem harder, we actually got a fundamentally different type of problem.”
Aparna Dhinakaran
“What if the best way to an evaluate an agent was actually with an agent.”
Aparna Dhinakaran
“LLM as a judge just gives you a fixed rubric with these fixed scores.”
Aparna Dhinakaran
videoWhy AI Agents Need Million-Token Context — Thomas Wolf & Olive Song, MiniMax
videoYour company brain will leak secrets: how we stopped it for big banks — Tanmai Gopal, PromptQL
videoTethered: Our Agents Are Us — Shu Fang, Two Sigma
videoAgents' next frontier: agent-to-agent and network effects — Jean-Denis Greze, TownChecking sign-in…
Loading comments…