Today we are open sourcing Ling-3.0-flash-dspark, a DSpark draft model built specifically for Ling-3.0-flash. On 4 NVIDIA Blackwell GPUs at batch 1, it delivered 1,120 tok/s, 0.78 ms mean TPOT, and an accept length of 9.95 across 1,000 requests. 🧵

Ling-3.0-flash has 42 layers: 35 KDA linear attention layers and 7 MLA full attention layers. Its MoE contains 512 routed experts and one shared expert, with top 8 plus 1 active per token. This hybrid design keeps attention cost low, even at 8K context.

Batch 1 exposes every microsecond. There is no larger batch to absorb launch cost or fill pipeline gaps. On the NEXTN path, our systems work raised mean output throughput from 288 to 606 tok/s and reduced mean TPOT from 3.33 to 1.53 ms.

At batch 1, mean TPOT is roughly step time divided by accept length. We worked on both. SGLang shortened each target step. DSpark uses a parallel draft backbone, a lightweight Markov head for local dependency, and a confidence head for verification scheduling.
Batch-1 latency is what a single-user actually feels. Open draft weights plus the scheduling fixes described here push mean TPOT to 0.78 ms with an accept length near 10 on four Blackwell GPUs.
Checking sign-in…
Loading comments…