← All IntelClip / AI ToolsFrom 'knowing' to 'doing': reliability as the autonomy bottleneck
From Data and Environment Curation for Post-Training LLMs — Mahesh Sathiamoorthy, Bespoke Labs · ≈2:37
“we have moved on from knowing to doing right so that's the idea of agents obviously”
“It's basically reliability, right?”
“at some point something falls apart like either they called the wrong tool or they made a mistake and you know what what not, right?”
What’s in it
- Traces how AI evaluation shifted from knowledge tests to agentic action benchmarks
- Names reliability as the core blocker to long-horizon agent autonomy
- Explains why agents fail on multi-step tasks (wrong tool calls, mistakes)
Clip transcript
we kind of think about the other thing I want to kind of talk about is you know how you know uh AI has evolved right. So early on we used to think about and evaluate models on what they know. Um for example this is a this was a very popular benchmark on uh testing LLMs on various kinds of STEM humanities and all that knowledge and these days we have all these benchmarks that test uh how how agents are able to do things we have moved on from knowing to doing right so that's the idea of agents obviously and the one of the key principles or one of the key things about agents is that they are autonomous And there are as I was saying there are many benchmarks including uh swb bench terminal bench and so on but ultimately for many people what they care about is are these agents autonomous for long durations of time uh Nick had uh sorry uh Ross had a great talk on long horizon right so that's the goal is eventually we make these agents autonomous for maybe few hours or you know few days or few weeks and what is it that's blocking the uh autonomy of agents. It's basically reliability, right? So, at some point something falls apart like either they called the wrong tool or they made a mistake and you know what what not, right? And what what's one lever to improve reliability there?
Comments
Sign in to comment.
Loading comments…