OpenAI's prompt cache makes a request 90% cheaper, but the cache key tops out around 15 requests per second. @HeggieConnor on how @unifygtm built its own routing around that limit, landing them close to a 95% cache hit rate.
@HeggieConnor @unifygtm Watch or listen to the latest Max Agency on your favorite podcasting platform. 🎧 Apple: https://t.co/VfE5h4auAT 🎧 Spotify: https://t.co/UhEMNqwxzf ⏯️ YouTube: https://t.co/Ldy6KnVu1O

Past roughly 15 requests per second on one cache key you quietly lose the caching discount. Deliberate routing across keys keeps hit rates near 95% at higher throughput.
Checking sign-in…
Loading comments…