Vibeleaderboard
← All Intel
Intel / video

Why 80% Reliability Isn't Good Enough — Felipe Blanes, Amazon AGI Lab

Source
youtube.com
Author
AI Engineer
Date
Why it matters

Shows how to build evals from real customer success criteria and signals instead of static benchmarks, and why reliability thresholds, not average scores, determine whether users trust a browser agent.

Key takeaways · AI-distilled
  • Blanes calls the core failure the " illusion": static look strong until customers do things nobody expected, and then teams are left hoping the holds up.
  • His eval flywheel has four steps: define success the way the customer does, capture signals (including by talking to customers), diagnose whether gaps sit in the model, or product, then feed findings into decisions.
  • He describes a "trust cliff": at 80% reliability an agent feels like more work, while around 92% it starts to feel trustworthy.
  • He argues that being open about an agent's limits builds more trust than higher benchmark scores.
Terms in this piece · Glossary
  • benchmark — A standard public test set for comparing AI models — the shared scoreboards behind every "model X beats model Y" claim.
  • eval — A repeatable test for AI quality — a set of tasks plus scoring — used the way software teams use test suites, because model output is too variable to judge by eyeballing.
  • AI agent — An AI system that doesn't just answer once but works toward a goal in a loop — taking actions, reading the results, and deciding what to do next.
  • agent harness — The scaffolding around a model that turns it into a working agent — the loop, the tools it can call, and the rules for when to stop.
Read the source www.youtube.com
More from AI Engineer
Recommended reads
Comments

Checking sign-in…

Loading comments…