
A dense, opinionated field guide to evals covering error analysis, design, and critique shadowing — the practical workflows for actually improving AI products, not just theory. Useful if you're building AI features and drowning in traces without a systematic way to measure or debug quality.
“In the projects we’ve worked on, we’ve spent 60-80% of our development time on error analysis and evaluation .”
Hamel Husain
“If you’re passing 100% of your evals, you’re likely not challenging your system enough. A 70% pass rate might indicate a more meaningful evaluation that’s actually stress-testing your application.”
“All you get from using these prefab evals is you don’t know what they actually do and in the best case they waste your time and in the worst case they create an illusion of confidence that is unjustified.”
“Use LLMs to scale what you’ve learned, not to avoid looking at data.”
Checking sign-in…
Loading comments…