← All IntelIntel / articleDo Automated Evals Work?
- Source
- hamel.dev
- Date
Why it matters
- Evidence on when LLM-judged evals agree with human annotation — and when they quietly diverge.
- We compared 100 human annotated traces against automated eval systems.
- Here's what we found.
Transcript
We compared 100 human annotated traces against automated eval systems. Here's what we found.
Comments
Sign in to comment.
Loading comments…