Why experts writing AI evals is creating the fastest-growing companies in history | Brendan Foody
Source
youtube.com
Author
Lenny's Podcast
Date
Why it matters
Explains that labs are bottlenecked on measuring success, so expert-written evalA repeatable test for AI quality — a set of tasks plus scoring — used the way software teams use test suites, because model output is too variable to judge by eyeballing.Full definition → drive model improvement. This clarifies where training signal for frontier models comes from.
Terms in this piece · Glossary
eval — A repeatable test for AI quality — a set of tasks plus scoring — used the way software teams use test suites, because model output is too variable to judge by eyeballing.