
If you're serving large models like Kimi or GLM yourself, this breaks down practical levers — and weight compression — for cutting GPU memory pressure without silently degrading output quality.
“On Kimi K2.6, that raises the amount of context we can hold in memory from roughly 686,000 tokens to about 1.37 million, twice as much.”
Cloudflare
“BF16 runs out of cache at 32 concurrent requests and can't admit a 33rd, while FP8 keeps going to 64 and reaches 2,192 tokens per second, about 41% higher than BF16's peak, for roughly 30% less cost per token.”
Cloudflare
“The checkpoint shrinks from 705 GB to 421 GB, about 40%, and per-GPU memory across an 8-way tensor-parallel deployment drops from roughly 88 GB to 52 GB, which leaves room for around 1.18 million tokens of KV cache on the same hardware.”
Cloudflare
“at our request volumes, even a one-in-a-billion mistake would show up regularly.”
Cloudflare
“The cost is under 1% on both throughput and tail latency, and even the upper bound of the 95% confidence interval stays near 1%.”
Cloudflare
Checking sign-in…
Loading comments…