Smaller, faster, safer: running Kimi and GLM at scale
- Source
- cfl.re
- Author
- Cloudflare
- Date

If you're serving large models like Kimi or GLM yourself, this breaks down practical levers — and weight compression — for cutting GPU memory pressure without silently degrading output quality.
- open weights — A model whose trained parameters are published for anyone to download and run — unlike API-only models you can access but never possess.
- KV cache — The memory a model keeps about text it has already read, so generating each new token doesn't require reprocessing the whole conversation.
- quantization — Shrinking a model by storing its numbers less precisely — like rounding — so it runs faster and fits on smaller hardware, at a small quality cost.
“On Kimi K2.6, that raises the amount of context we can hold in memory from roughly 686,000 tokens to about 1.37 million, twice as much.”
Cloudflare
“BF16 runs out of cache at 32 concurrent requests and can't admit a 33rd, while FP8 keeps going to 64 and reaches 2,192 tokens per second, about 41% higher than BF16's peak, for roughly 30% less cost per token.”
Cloudflare
“The checkpoint shrinks from 705 GB to 421 GB, about 40%, and per-GPU memory across an 8-way tensor-parallel deployment drops from roughly 88 GB to 52 GB, which leaves room for around 1.18 million tokens of KV cache on the same hardware.”
Cloudflare
“at our request volumes, even a one-in-a-billion mistake would show up regularly.”
Cloudflare
“The cost is under 1% on both throughput and tail latency, and even the upper bound of the 95% confidence interval stays near 1%.”
Cloudflare
postCloudflare AI Search goes GA with image embeddings and PDF OCR
articleWe are introducing Clef and Clef-flash, open-source decision models hosted on…
articleToday, we’re making the Cloudflare Monetization Gateway available as part of a…
articleCloudflare Auto Router picks cheaper models per request in AI Gateway
Checking sign-in…
Loading comments…

