Vibeleaderboard
← All Intel
Intel / post

Ling-3.0-flash gets INT4 and MXFP4 builds that fit one DGX Spark

Source
Ant Ling
Date
Ant Ling@AntLingAGI
Thread · 5 parts

🚀 Today, we’re releasing INT4 and FP4 (MXFP4) variants of Ling-3.0-flash. Both run end to end on a single NVIDIA DGX Spark via our Spark-adapted SGLang path. For FP4, W4A16 is the stable default, while W4A8 is tuned for higher throughput. The efficiency and accuracy of the quantized models are still evolving. We will continue updating the inference implementation and quantized weights over time.

🎬 Demo: Ling-3.0-flash MXFP4 running locally on one DGX Spark. In our tests: ⚡ ~80 tok/s decoding 📥 2,500–3,500 tok/s long-input prefilling 👥 Smooth use by 3–4 concurrent users 🔒 Private, on-device inference for coding, agents, and offline batch jobs. 👇

🩺 Demo: A clinical data pipeline powered by Ling-3.0-flash, running entirely on one DGX Spark. 🔒 Local de-identification keeps patient data private 🤖 Extracts vitals, lab results and ultrasound findings, then assigns study groups 🚩 Flags ambiguous cases for human review

Download the weights: 🤗 Hugging Face: FP4: https://t.co/hRGMeO7sUo INT4: https://t.co/UBzvHyVVnm 🤖 ModelScope: FP4: https://t.co/bfhiJSlGQM INT4: https://t.co/LnJICZWmnV

Read the full thread on X

Context

Ant Ling releases INT4 and MXFP4 versions of Ling-3.0-flash, its 124-billion-parameter -focused model, sized to run end to end on a single NVIDIA DGX Spark desktop unit through what the company calls a Spark-adapted SGLang serving path. For the FP4 build, the company says W4A16 is the stable default configuration while W4A8 is tuned instead for higher throughput. The release follows the model's BF16 and FP8 weights, which shipped the day before.

In its own testing, Ant Ling reports about 80 tokens per second decoding and 2,500 to 3,500 tokens per second for long-input prefilling on the MXFP4 build, with smooth use by three to four concurrent users on one Spark unit. The company frames this as enabling private, on-device for coding, agents and offline batch jobs, and shows a clinical-data demo that locally de-identifies and extracts information from medical records while flagging ambiguous cases for human review; that demo was presented as video and is described here only as the company's written summary, not independently verified. Ant Ling says the weights and inference implementation are still being updated, so current accuracy and throughput figures may not be final.

Terms in this piece · Glossary
  • AI agent — An AI system that doesn't just answer once but works toward a goal in a loop — taking actions, reading the results, and deciding what to do next.
  • inference — Running a trained model to get answers — the phase where AI is actually used, as opposed to trained.
  • quantization — Shrinking a model by storing its numbers less precisely — like rounding — so it runs faster and fits on smaller hardware, at a small quality cost.
  • mixture-of-experts — A model built from many specialist sub-networks where only a few activate per token, giving big-model capability at small-model running cost.
More from Ant Ling
Recommended reads
Comments

Checking sign-in…

Loading comments…