Stanford CS329A Self-Improving AI Agents | Part 8 | Agentic Evaluations and Long Horizon Tasks
Source
youtube.com
Author
Stanford Online
Date
Why it matters
Explains why agentic evalA repeatable test for AI quality — a set of tasks plus scoring — used the way software teams use test suites, because model output is too variable to judge by eyeballing.Full definition → are harder than prior benchmarks and how task time horizon is used to measure long-task capability. Helpful for designing your own evals.
Terms in this piece · Glossary
eval — A repeatable test for AI quality — a set of tasks plus scoring — used the way software teams use test suites, because model output is too variable to judge by eyeballing.