Vibeleaderboard
← All Intel
Intel / article

Better prompt caching for GPT-6

Source
openai.com
Date
Terms in this piece · Glossary
  • context windowThe maximum amount of text a model can consider at once — its working memory for the current conversation or task.
Why it matters

Higher cache hit rates and explicit breakpoints directly reduce latency and per-request cost for agents that repeatedly send large, mostly-unchanged to GPT-6.

Recommended reads
Comments

Checking sign-in…

Loading comments…