
It gives you empirical production numbers — not vendor benchmarks — on which models are actually absorbing volume versus where the money goes, useful when deciding whether to route cheap models for bulk work and reserve premium models for the paths that need them. The per-modality breakdown also shows what real teams are picking for image and video generation.
“Open-weight models ran 29% of gateway tokens in June on just under 4% of spend. That is nearly a third of the tokens for one twenty-fifth of the dollars.”
“This is evidence of the routing discipline June's report documented, now visible in the aggregate: high-volume work goes to low-cost models, high-risk work stays on the frontier.”
“Roughly one in eight enterprise customers now runs an open-weight model in production.”
“And on current trajectories, an open-weight lab will soon be the second-largest by volume on AI Gateway.”
“The market spent more overall, but not more per token.”
Checking sign-in…
Loading comments…