Vibeleaderboard
← All Intel
Intel / video

Build Hour: Prompt Caching

Source
youtube.com
Author
OpenAI
Date
Why it matters

Explains how caching works and how to structure prompts to hit it, a direct way to reduce latency and cost in and API workloads.

Terms in this piece · Glossary
  • prompt caching — Reusing the model's processed form of a repeated prompt prefix so subsequent calls skip re-reading it, cutting cost and latency substantially.
  • token — The chunk of text a model reads and writes in — roughly three-quarters of a word — and the unit AI usage is billed in.
  • AI agent — An AI system that doesn't just answer once but works toward a goal in a loop — taking actions, reading the results, and deciding what to do next.
Read the source www.youtube.com
More from OpenAI
Recommended reads
Comments

Checking sign-in…

Loading comments…