← All IntelClip / OtherSmarter models are more token-efficient with tools
From The State of Model Routing — NVIDIA, Cognition, OpenRouter · ≈23:47
“It's not obvious that actually more expensive models are actually creating an overall cheaper system.”
“like the scaling laws uh if you have a larger model, it's going to be more efficient with its tokens.”
What’s in it
- Explains why bigger AI models end up cheaper to run per task
- Reveals a trick: sub-agents reference files instead of dumping full text
- Shows how smarter models self-select which files or commands actually matter
Clip transcript
cheaper and less of a problem in future? >> Yeah, it's definitely possible as well. I've seen cases where um you know, the psychic agent does a bunch of work. It tells the main model, "Oh yeah, like here's here's all the things I found." Instead of dumping the full thing, it just references them by file. And then the main model is actually generally you find these larger, smarter models, they're actually more token efficient with how they use tools and how they read. And so, they actually read the files in a way where they only see the important parts, right? Or they decide that, "Oh, actually I only need to look at a subset of this." Or, "Oh, I can run a single command and just know if everything is done properly." Um, it's actually quite amazing um the fact that you know, these multi-model systems they actually seem to scale and get better with intelligence, which is um not something we should just take for granted, right? It's not obvious that actually more expensive models are actually creating an overall cheaper system. >> Yeah, like the scaling laws uh if you have a larger model, it's going to be more efficient with its tokens. Uh smaller models less efficient with its tokens.
Comments
Sign in to comment.
Loading comments…