← All IntelClip / AI AgentsCapability is solved; reliability has to be in the nines
From Perception Agents — Antje Barth, Amazon AGI Lab · ≈3:33
Reframes 60-80% task success as unusable, giving a concrete bar for when an agent can actually be trusted with real work.
What’s in it
- Reframes 60-80% task success as unusable, giving a concrete bar for when an agent can actually be trusted with real work.
Clip transcript
capabilities to models. Now the next hard part is really reliability and without reliability we cannot really build up trust in those systems. So here's a quick gut check and maybe all of you can just think about an agent doing work in an end toend workflow. How often do you think that actually succeeds these days? Maybe 60 maybe 80% of the time. And it sounds really fine, but if you look into this, if your agent one in four times deletes a database, you will never touch that agent again, right? So when you need this reliability, you really need to be it in the nines. You need to have the trust that it actually can do the work successfully.
Comments
Sign in to comment.
Loading comments…