← All IntelClip / AI AgentsEval usage stats at scale
From The Future of Evals: From LLM as a Judge to Agent as a Judge — Aparna Dhinakaran, Arize AI · ≈1:09
“We run over 100 million evals every month.”
“The average team runs about 12 different eval jobs with the top teams running over 3,800 different evaluators.”
“This is actually what's helping teams figure out what's working, catch their failures, and that's the type of data you need to fuel your continual learning”
What’s in it
- Reveals real usage stats from a platform running 100M+ evals monthly
- Shows why top teams run thousands of evaluators, not just dozens
- Explains how production trace evals fuel continual model improvement
Clip transcript
running on their live production agent via their traces. Little bit of some stats for you guys. We run over 100 million evals every month. The average team runs about 12 different eval jobs with the top teams running over 3,800 different evaluators. And offline evals, online evals, they each have their own place, but today what I'm actually going to talk to you about is the teams that are running evals on their traces. This is actually what's helping teams figure out what's working, catch their failures, and that's the type of data you need to fuel your continual learning
Comments
Sign in to comment.
Loading comments…