← All IntelClip / AI AgentsWhy human-hour benchmarks break down
From Rethinking Environments for Long-Horizon Work — Rayan Garg, Theta Software · ≈4:37
“Um, you know, what's long horizon for a human isn't necessarily that difficult for a model depending on what the actual task you care about is.”
AI Engineer
“It might take them like days to do that if it's a really big Excel file, but for a model, it can maybe write a Python script or find some other cool trick to do that really quickly.”
AI Engineer
“And if you, you know, someone is out there saying, "Hey, we have some tasks or environments that are 16 hours long on average, someone else is 20 hours." That's really hard to compare across people because there's so many different things in the methodology that really impact uh kind of what that actually means.”
AI Engineer
“It can mean, you know, the quality of the experts you're using. Some more experienced experts might actually be way more efficient at doing a certain type of financial or coding task, whatever it kind of is.”
AI Engineer
What’s in it
- Concrete postmortem on benchmark methodology: tedious-for-human tasks can be trivial for models (Excel reformatting via a script), and human-hour estimates get very noisy at the top-1% skill frontier, so cross-vendor hour claims aren't comparable.
Clip transcript
and steps, but there's also a lot of weaknesses in the other approach of kind of relying on humans. Um, you know, what's long horizon for a human isn't necessarily that difficult for a model depending on what the actual task you care about is. You know, there's a lot of tasks that are really tedious and time inensive. Maybe like, you know, some financial analyst has to go into like an Excel file and fix a bunch of formatting issues throughout uh the the task. Maybe they're changing like the theming of like the colors in the in the actual file, right? That might be really tedious for a human to take. It might take them like days to do that if it's a really big Excel file, but for a model, it can maybe write a Python script or find some other cool trick to do that really quickly. And that's not really hard for it to do, but you never really expect a financial expert to do that because, you know, they most of them don't really know how to write these Python scripts. Um, so I think that's one thing to note. And the other that I kind of briefly touched upon before is that the methodology of how we actually measure this has a really big impact. And if you, you know, someone is out there saying, "Hey, we have some tasks or environments that are 16 hours long on average, someone else is 20 hours." That's really hard to compare across people because there's so many different things in the methodology that really impact uh kind of what that actually means. It can mean, you know, the quality of the experts you're using. Some more experienced experts might actually be way more efficient at doing a certain type of financial or coding task, whatever it kind of is. Um and I think this becomes really really important as we start shifting towards uh kind of the frontier of even human capabilities. So you know as this meter talks about this but as you shift towards more long resin tasks and tasks that only the top 10% the top 1% top.1% of humans can really do these estimates start to get really really noisy and
Comments
Checking sign-in…
Loading comments…