Redundant renderings of the same function eat context windowThe maximum amount of text a model can consider at once — its working memory for the current conversation or task.Full definition → slots useful code needed. Organizing retrieved evidence by canonical code object, keeping a companion fragment only when it adds semantics, took Recall@4K from 39.6% to 63.7% on 36% fewer evidence tokenThe chunk of text a model reads and writes in — roughly three-quarters of a word — and the unit AI usage is billed in.Full definition →.
Terms in this piece · Glossary
context window — The maximum amount of text a model can consider at once — its working memory for the current conversation or task.
RAG — Retrieval-augmented generation — fetching relevant documents first and pasting them into the model's context so it answers from your data instead of memory.
token — The chunk of text a model reads and writes in — roughly three-quarters of a word — and the unit AI usage is billed in.