Vibeleaderboard
← All Intel
Intel / article

Smaller, faster, safer: running Kimi and GLM at scale

Source
cfl.re
Author
Cloudflare
Date
Why it matters

If you're serving large models like Kimi or GLM yourself, this breaks down practical levers — and weight compression — for cutting GPU memory pressure without silently degrading output quality.

Terms in this piece · Glossary
  • open weights — A model whose trained parameters are published for anyone to download and run — unlike API-only models you can access but never possess.
  • KV cache — The memory a model keeps about text it has already read, so generating each new token doesn't require reprocessing the whole conversation.
  • quantization — Shrinking a model by storing its numbers less precisely — like rounding — so it runs faster and fits on smaller hardware, at a small quality cost.
Key quotes

“On Kimi K2.6, that raises the amount of context we can hold in memory from roughly 686,000 tokens to about 1.37 million, twice as much.”

Cloudflare

“BF16 runs out of cache at 32 concurrent requests and can't admit a 33rd, while FP8 keeps going to 64 and reaches 2,192 tokens per second, about 41% higher than BF16's peak, for roughly 30% less cost per token.”

Cloudflare

“The checkpoint shrinks from 705 GB to 421 GB, about 40%, and per-GPU memory across an 8-way tensor-parallel deployment drops from roughly 88 GB to 52 GB, which leaves room for around 1.18 million tokens of KV cache on the same hardware.”

Cloudflare

“at our request volumes, even a one-in-a-billion mistake would show up regularly.”

Cloudflare

“The cost is under 1% on both throughput and tail latency, and even the upper bound of the 95% confidence interval stays near 1%.”

Cloudflare
More from Cloudflare
Recommended reads
Comments

Checking sign-in…

Loading comments…