Vibeleaderboard
← All Intel
Intel / blog

Cut your AI spend with AI Gateway's Auto Router

Source
blog.cloudflare.com
Date
Key takeaways · AI-distilled
  • On Cloudflare's internal knowledge-work (97 tasks, three samples each), cloudflare/auto succeeded 86.6% of the time for $2.10, versus 96.6% at $5.91 for Claude Opus 5.5 and 84.2% at $2.64 for GPT-6 Sol. Cost per success: $0.0084, $0.0210, $0.0108.
  • The router first filters models by request format, credentials, access policies, spend limits and provider health. A Workers AI classifier then reads recent turns, scores 14 task categories, and rates complexity, ambiguity, stakes and dependence from 1 to 5.
  • A scoring matrix combines those signals with model benchmarks and prices. Price weighs more on easy requests and the cost penalty shrinks as difficulty rises. Adding a new model needs only its benchmark-derived weights, not retraining.
  • Switching models discards the prompt cache, so the router avoids switching within a turn and applies a cross-turn penalty that grows with tokens in context, pricing the current model at its cache-read rate and every other candidate at the full cost of rewriting context.
  • Cloudflare argues cheaper per-token models do not always give cheaper outcomes, since some use far more tokens, so a router should minimize predicted trajectory cost. A planned cloudflare/auto-best will pick the highest expected quality without the cost tradeoff.
Terms in this piece · Glossary
  • benchmark — A standard public test set for comparing AI models — the shared scoreboards behind every "model X beats model Y" claim.
  • context window — The maximum amount of text a model can consider at once — its working memory for the current conversation or task.
  • token — The chunk of text a model reads and writes in — roughly three-quarters of a word — and the unit AI usage is billed in.
Why it matters

Lets teams cut frontier-model spend across Claude Code, Codex and OpenCode harnesses without asking users to choose models. Cloudflare reports up to 30% savings internally.

Read the source blog.cloudflare.com
Recommended reads
Comments

Checking sign-in…

Loading comments…