
Gives you precise, shared terminology for building harnesses and explains specifically why agentic evals fail differently than single-turn evals (compounding errors, valid-but-unexpected solutions), which helps you design graders and suites that don't break on production agents.
“These same capabilities that make AI agents useful—autonomy, intelligence, and flexibility—also make them harder to evaluate.”
“Without them, it’s easy to get stuck in reactive loops—catching issues only in production, where fixing one failure creates others.”
“It “failed” the evaluation as written, but actually came up with a better solution for the user.”
“Agents use tools across many turns, modifying state in the environment and adapting as they go—which means mistakes can propagate and compound.”
Checking sign-in…
Loading comments…