Vibeleaderboard
← All Intel
Intel / post

Keeping Transformed Weights Resident Cuts SGLang Cold Starts to Seconds

Source
Ant Ling
Date
Ant Ling@AntLingAGI
Thread · 10 parts

Today, we’re introducing the Weight Cache Daemon for SGLang. 🚀 On Ling-2.6-1T FP8, it reduced weight loading to ~0.63s, up to ~780× faster than disk loading, and cut total engine startup from 8.8 minutes to ~0.53 minutes. Here’s how it works.

What made cold starts so slow? 🔍 Weight loading dominated the startup path. Each TP rank had to read around 120 GB from disk, deserialize it, apply TP sharding, run FP8 quantization, and repack the weights. Every restart repeated the same deterministic work.

The solution is a persistent GPU process that keeps transformed weights ready for reuse. One daemon runs per GPU and TP rank. It loads and transforms weights once, keeps them in GPU memory, and serves CUDA IPC handles over a Unix socket.

On restart, the engine initializes on the meta device, with no CPU or GPU storage allocated. It points each https://t.co/lYulEBoKPh directly to an IPC-mapped tensor, including FP8 weight_scale and other post-quantization tensors. No copy. No disk I/O. No requantization.

Read the full thread on X

Context

Restarting a large server repeats the same expensive setup: each tensor-parallel worker reads about 120GB of weights from disk, shards them and runs FP8 . Ant Ling describes a Weight Cache Daemon for SGLang, developed with Alibaba and the SGLang team, in which one persistent process per GPU and rank loads and transforms weights once, keeps them in GPU memory, and shares them with new engine processes through CUDA IPC handles over a Unix socket. A restarting engine initializes on the meta device with no storage allocated and points its tensors at the shared memory, so, the thread says, there is no copy, disk read or requantization.

On Ling-2.6-1T FP8 the company reports weight loading of about 0.63 seconds and total engine startup falling from 8.8 minutes to about 0.53 minutes. It also describes an active-standby setup where a backup engine can take over in under one second. These are Ant Ling's own figures. The cache is currently verified only for unquantized and block-wise FP8 models.

Terms in this piece · Glossary
  • inference — Running a trained model to get answers — the phase where AI is actually used, as opposed to trained.
  • quantization — Shrinking a model by storing its numbers less precisely — like rounding — so it runs faster and fits on smaller hardware, at a small quality cost.
More from Ant Ling
Recommended reads
Comments

Checking sign-in…

Loading comments…