← All IntelIntel / article
Better prompt caching for GPT-6
- Source
- openai.com
- Date
Terms in this piece · Glossary
- context window — The maximum amount of text a model can consider at once — its working memory for the current conversation or task.
Why it matters
Higher cache hit rates and explicit breakpoints directly reduce latency and per-request cost for agents that repeatedly send large, mostly-unchanged to GPT-6.
Read the source openai.com
Recommended reads
Comments
Checking sign-in…
Loading comments…
