model routing — Sending each request to a model chosen by the difficulty of the task, rather than using one model for everything.
token — The chunk of text a model reads and writes in — roughly three-quarters of a word — and the unit AI usage is billed in.
Why it matters
You can now see model routingSending each request to a model chosen by the difficulty of the task, rather than using one model for everything.Full definition →, fallback attempts, tokenThe chunk of text a model reads and writes in — roughly three-quarters of a word — and the unit AI usage is billed in.Full definition → cost and time-to-first-token per request inside your existing observability stack, with sampling and no prompt content leaving the gateway.