Vibeleaderboard
← All Intel
Intel / video

Why Tejal Patwardhan stopped underestimating the models - Episode 21

Source
youtube.com
Author
OpenAI
Date
Why it matters

Explains why saturated benchmarks fail and how OpenAI builds real-work with human baselines, useful for anyone designing evals for agents.

Terms in this piece · Glossary
  • eval — A repeatable test for AI quality — a set of tasks plus scoring — used the way software teams use test suites, because model output is too variable to judge by eyeballing.
  • context window — The maximum amount of text a model can consider at once — its working memory for the current conversation or task.
Read the source www.youtube.com
More from OpenAI
Recommended reads
Comments

Checking sign-in…

Loading comments…