Vibeleaderboard
← All Intel
Intel / video

Stanford CS329A Self-Improving AI Agents | Part 8 | Agentic Evaluations and Long Horizon Tasks

Source
youtube.com
Author
Stanford Online
Date
Why it matters

Explains why agentic are harder than prior benchmarks and how task time horizon is used to measure long-task capability. Helpful for designing your own evals.

Terms in this piece · Glossary
  • eval — A repeatable test for AI quality — a set of tasks plus scoring — used the way software teams use test suites, because model output is too variable to judge by eyeballing.
Read the source www.youtube.com
More from Stanford Online
Recommended reads
Comments

Checking sign-in…

Loading comments…