← All IntelClip / OtherOrchestration as a necessary controller layer
From The State of Model Routing — NVIDIA, Cognition, OpenRouter · ≈45:04
“I think applications, especially as built on non-deterministic systems like models, operate in a very low-trust environment.”
“if you look at like the the the rankings on OpenRouter, if you look at our our public data and you look at like the top model being used by dollar spent on classification tasks, well, guess what it is. It's Opus.”
What’s in it
- Explains why multi-agent AI systems need a central orchestration/arbitration layer
- Reveals surprising usage data: Opus dominates spend even on classification tasks
- Argues caching—not raw model power—will keep small models relevant long-term
Clip transcript
>> Yeah. Makes sense. >> I think applications, especially as built on non-deterministic systems like models, operate in a very low-trust environment. So yes, most of the improvements will likely be distributed across both models and the harnesses, but I think overall it's it's it's mostly it's There There will There will have to be some form of controller trying to have some form of arbitration because even from the model perspective you aren't in a perfectly visible world. You don't know the behavior of every model, so it's it's going to be at the orchestration level where you have these kind of things. And this has traditionally been shown by other industries like when web when web launched, you know, you had traffic-based routing. So it's different. But all the sort of routing controls have been centralized over time. >> Makes sense. >> I I it's most likely going to be good news in the future. Um and and I I think like caching is a big reason for that. Even if you I think like a to take the flip side of of this argument, um the you know, it might be that in the future we have like one big model that's like, "I know I am the like most efficient at everything and I'm like way more efficient than Haiku. I'll solve every task better than Haiku can at like a lower price. Um why should I ever delegate to Haiku?" There's something like that actually could could be a model that we have in the future. Um but you're always going to have these like, you know, for example, caching. It could be that like you tell the model that this other model like does have the right context in cash and uh you know, the the orchestrator model just always has more context and the models have to be aligned. So, I I think like it's I I don't really see a world where like we wouldn't be able to get models to collaborate really well and and I think they're going to get better over time. Um in part because they're you know, they just have limited memory. So, I I think that's kind of one one deciding factor and another is that um there will be like uh there will continue to be like if you just look at like the the the rankings on OpenRouter, if you look at our our public data and you look at like the top model being used by dollar spent on classification tasks, well, guess what it is. It's Opus. >> [laughter] >> I think there are there are there are big opportunities for like using small models for in distribution easy tasks and the and like as time goes on, that's going to be a larger and larger percentage of tasks relative to like the most valuable tasks that um very smart models spend most of their time on.
Comments
Sign in to comment.
Loading comments…