Prompt Caching Explained: Stop Overpaying for AI Agents
Source
youtube.com
Author
Hugging Face
Date
Why it matters
Prompt caching works differently from database caching, and a harness that breaks the prefix pays full price on every turn. Learn how to structure context to keep cache hits.
Key takeaways · AI-distilled
prompt cachingReusing the model's processed form of a repeated prompt prefix so subsequent calls skip re-reading it, cutting cost and latency substantially.Full definition → stores the processed input prefix, not the model's answer. Each agent turn resends the whole transcript plus the new message, so cached re-reads are billed at a discount, which the speaker puts at about 10% of the normal input price.
Cache lifetime varies by provider. The speaker says OpenAI keeps it for about an hour, while Anthropic defaults to roughly five minutes on the API. When he took long meal breaks mid-session, the cache expired and the next turn was billed at full price again.
Not every provider caches automatically. The speaker says OpenAI and Hugging Face inference providers do it by default, but with Anthropic or Gemini the agent has to turn caching on in its API calls. Chat-completions APIs often lack the default that Responses-style APIs have.
Anything dynamic in the system promptThe standing instructions a model receives before any user input — defining its role, rules, tools, and tone for the whole conversation.Full definition →, like a timestamp, the current working directory, or a tool list that changes, invalidates the cache for everything after it. Keep the history append-only. context compactionSummarizing an agent's earlier conversation to free room in the context window so a long session can keep going.Full definition → also resets the cache, which is expected but costs a full-price turn.
In the speaker's own session with DeepSeek V4 Flash through Hugging Face inference providers, about 10.9 million tokens moved through a roughly 126K context windowThe maximum amount of text a model can consider at once — its working memory for the current conversation or task.Full definition → for about 6 cents. He recommends a harness that shows cache hit rate per request and per session.
Terms in this piece · Glossary
prompt caching — Reusing the model's processed form of a repeated prompt prefix so subsequent calls skip re-reading it, cutting cost and latency substantially.
context window — The maximum amount of text a model can consider at once — its working memory for the current conversation or task.
system prompt — The standing instructions a model receives before any user input — defining its role, rules, tools, and tone for the whole conversation.
context compaction — Summarizing an agent's earlier conversation to free room in the context window so a long session can keep going.