context window — The maximum amount of text a model can consider at once — its working memory for the current conversation or task.
model routing — Sending each request to a model chosen by the difficulty of the task, rather than using one model for everything.
token — The chunk of text a model reads and writes in — roughly three-quarters of a word — and the unit AI usage is billed in.
Why it matters
A worked example of model routingSending each request to a model chosen by the difficulty of the task, rather than using one model for everything.Full definition → for coding agents: aggregate cost fell 57.1% at equal pass count, while tokenThe chunk of text a model reads and writes in — roughly three-quarters of a word — and the unit AI usage is billed in.Full definition → use rose 3.5x and some tasks regressed.