Why Tejal Patwardhan stopped underestimating the models - Episode 21
Source
youtube.com
Author
OpenAI
Date
Why it matters
Explains why saturated benchmarks fail and how OpenAI builds real-work evalA repeatable test for AI quality — a set of tasks plus scoring — used the way software teams use test suites, because model output is too variable to judge by eyeballing.Full definition → with human baselines, useful context windowThe maximum amount of text a model can consider at once — its working memory for the current conversation or task.Full definition → for anyone designing evals for agents.
Terms in this piece · Glossary
eval — A repeatable test for AI quality — a set of tasks plus scoring — used the way software teams use test suites, because model output is too variable to judge by eyeballing.
context window — The maximum amount of text a model can consider at once — its working memory for the current conversation or task.