Evals Skills for Coding Agents
hamel.dev- Category
- Other
- Type
- ARTICLE
- Builder
- @HamelHusain
- Added
- Jul 21, 2026
About
Today, I’m publishing evals-skills , a set of skills for AI product evals 1 . They guard against common mistakes I’ve seen helping 50+ companies and teaching 4,000+ students in our course . Why Skills for Evals Coding agents now instrument applications, run experiments, analyze data, and build interfaces. I’ve been pointing them at evals. OpenAI’s Harness Engineering article makes the case well. They built a product entirely with Codex agents — three engineers, five months, ~1 million lines of c
Why it made the leaderboard
If your coding agent is instrumenting and evaluating your AI product, these skills encode the error-analysis discipline that keeps it from lumping distinct failure modes into one useless score.
Media

Comments (0)
No comments yet
Indexed by a proprietary survey. Corrections welcome.