Vibeleaderboard
← All Intel
Intel / article

Claude’s new auto eval tool

Source
Hamel Husain
Author
Hamel Husain
Date
Key takeaways · AI-distilled
  • Hamel Husain and Isaac Flath found the build_ workflow in Anthropic's claude-api plugin asked them to pick a failure to turn into an eval before they had reviewed any conversations. Husain argues you should look at the data first, then decide which failures matter.
  • To validate labels, the tool had reviewers skim long conversations in a Markdown file and report corrections in chat. The pair asked Claude to build a web annotation app instead, and Husain argues most eval work belongs in a web app rather than a chat window.
  • The generated call-transfer evaluator bundled four checks, one and three code-based, into a single pass/fail score. Husain recommends scoping each eval to one error, or at least separating code checks from judge checks, and reading the judge prompt itself.
  • On the positive side, Husain says the plugin's one-shot issue discovery surfaced problems with human handoff, formatting and voice behavior that other auto-eval approaches had missed, the strongest result of that kind he has seen.
  • His verdict is to hold off for now. The plugin's author told him it would be updated in response to the feedback, so the specific workflow reviewed here may change.
Terms in this piece · Glossary
  • eval — A repeatable test for AI quality — a set of tasks plus scoring — used the way software teams use test suites, because model output is too variable to judge by eyeballing.
  • context window — The maximum amount of text a model can consider at once — its working memory for the current conversation or task.
  • LLM-as-judge — Using one model to score another's output against a rubric, so quality can be measured at a scale human grading cannot reach.
Why it matters

Anthropic's eval tooling will shape how many teams build evals. This review shows where it nudges you to skip error analysis and validate judgments without enough .

More from Hamel Husain
Recommended reads
Comments

Checking sign-in…

Loading comments…