Teams bolting memory onto agents usually assume more retrieved context windowThe maximum amount of text a model can consider at once — its working memory for the current conversation or task.Full definition → is strictly better, and this shows accuracy and cost both turn on the dose and the delivery mode.
Terms in this piece · Glossary
distillation — Training a small, cheap model to imitate a big one's outputs, keeping much of the capability at a fraction of the cost.
context window — The maximum amount of text a model can consider at once — its working memory for the current conversation or task.