Why 80% Reliability Isn't Good Enough — Felipe Blanes, Amazon AGI Lab
Source
youtube.com
Author
AI Engineer
Date
Why it matters
Shows how to build evals from real customer success criteria and signals instead of static benchmarks, and why reliability thresholds, not average scores, determine whether users trust a browser agent.
Key takeaways · AI-distilled
Blanes calls the core failure the "benchmarkA standard public test set for comparing AI models — the shared scoreboards behind every "model X beats model Y" claim.Full definition → illusion": static evalA repeatable test for AI quality — a set of tasks plus scoring — used the way software teams use test suites, because model output is too variable to judge by eyeballing.Full definition → look strong until customers do things nobody expected, and then teams are left hoping the AI agentAn AI system that doesn't just answer once but works toward a goal in a loop — taking actions, reading the results, and deciding what to do next.Full definition → holds up.
His eval flywheel has four steps: define success the way the customer does, capture signals (including by talking to customers), diagnose whether gaps sit in the model, agent harnessThe scaffolding around a model that turns it into a working agent — the loop, the tools it can call, and the rules for when to stop.Full definition → or product, then feed findings into decisions.
He describes a "trust cliff": at 80% reliability an agent feels like more work, while around 92% it starts to feel trustworthy.
He argues that being open about an agent's limits builds more trust than higher benchmark scores.
Terms in this piece · Glossary
benchmark — A standard public test set for comparing AI models — the shared scoreboards behind every "model X beats model Y" claim.
eval — A repeatable test for AI quality — a set of tasks plus scoring — used the way software teams use test suites, because model output is too variable to judge by eyeballing.
AI agent — An AI system that doesn't just answer once but works toward a goal in a loop — taking actions, reading the results, and deciding what to do next.
agent harness — The scaffolding around a model that turns it into a working agent — the loop, the tools it can call, and the rules for when to stop.