Automating eval design and hillclimbing with Claude Code
Source
ClaudeDevs
Author
ClaudeDevs
Date
Terms in this piece · Glossary
eval — A repeatable test for AI quality — a set of tasks plus scoring — used the way software teams use test suites, because model output is too variable to judge by eyeballing.
agent skill — A reusable instruction file that teaches an agent how to do one job well — the procedure, the tools, and what counts as done.
Why it matters
Shows how to design evalA repeatable test for AI quality — a set of tasks plus scoring — used the way software teams use test suites, because model output is too variable to judge by eyeballing.Full definition → that avoid fooling yourself and hillclimb one change at a time with a held-out set. The build-eval and hillclimb commands automate this inside Claude Code.