Your 5-minute cache is already dead when you sit down. API models don't keep the thread. Every turn resends the whole history. A one-token tool result still pays that whole history. On Opus 5, 1 million uncached tokens are $5. A hit is $0.50. Cache-hit discounts are provider-specific. > Anthropic 0.1x on a hit > OpenRouter blend near 1/5 > Grok 4.6 cache read 0.25x > DeepSeek Flash off-peak 0.03x Earlier, we ran 4.5 through 12 turns at 30,000 tokens. Turns 2 and 10 were full-price re-reads. Compact before you walk away so the miss stays small. Buy the 1-hour write at 2x, and a return can hit. Claude Code and Codex often compact around 200,000 to 300,000 tokens. Filling 1 million on purpose buys the expensive part. So, What do you pin so the cheap copy stays cheap? Where does the project live so compact doesn't wipe it? Why is dollars per million the wrong scoreboard? Full Breakdown ↓↓


Long sessions quietly pay full price to re-read history once the five-minute cache lapses. Knowing your provider's hit multiplier, when to buy the 2x one-hour write, and when to compact turns that into a controllable cost.
Checking sign-in…
Loading comments…