← All IntelClip / OtherOrchestrator evaluation could now be 90% faster with agentic workflows
From Benchmarking Coding Agents on New vs Legacy Codebases — Denys Linkov, Wisedocs · ≈3:47
“So we spent around 2 months evaluating orchestrators for our AI pipeline.”
“Now, I'm pretty confident we could do this 90% faster now with the tooling that we have.”
“It's very easy to undergo AI psychosis, where you look at a deep research report that's 20 pages long and you say, "Wow, this looks good."”
What’s in it
- Reveals a real-world process for evaluating 5 open-source AI orchestrators
- Shows how agentic workflows could cut evaluation time by 90%
- Warns about 'AI psychosis' — trusting flashy AI reports over reality
Clip transcript
out. So we spent around 2 months evaluating orchestrators for our AI pipeline. We looked at five open-source projects and we wanted to benchmark and see how effective they were for our use case. And we started this off before deep research came out as part of Google and OpenAI, so that web search capability to do a comprehensive analysis was still not there. Now, after we actually gathered these requirements, we built out proof of concepts with a team of three to make sure that we actually got the right results. Now, I'm pretty confident we could do this 90% faster now with the tooling that we have. Before we would manually go through, use a little bit of AI, but put everything into a Confluence doc, and we'd evaluate across 17 different criteria that we came up with. Nowadays, we could build a much more agentic workflow to do that, starting off with deep research, uh making sure that we match that against the problem statements that we have, creating sub-agents for each of these criteria and uh products, and then finally building POCs and evaluating. So, things have changed in the past year and a half where we could actually go much, much faster. But, we still have to maintain that same set of quality because it's very easy to undergo AI psychosis, where you look at a deep research report that's 20 pages long and you say, "Wow, this looks good." And then those features don't actually exist in the product, and you've set yourself back.
Comments
Sign in to comment.
Loading comments…