🚀 Today, we’re releasing INT4 and FP4 (MXFP4) variants of Ling-3.0-flash. Both run end to end on a single NVIDIA DGX Spark via our Spark-adapted SGLang path. For FP4, W4A16 is the stable default, while W4A8 is tuned for higher throughput. The efficiency and accuracy of the quantized models are still evolving. We will continue updating the inference implementation and quantized weights over time.
🎬 Demo: Ling-3.0-flash MXFP4 running locally on one DGX Spark. In our tests: ⚡ ~80 tok/s decoding 📥 2,500–3,500 tok/s long-input prefilling 👥 Smooth use by 3–4 concurrent users 🔒 Private, on-device inference for coding, agents, and offline batch jobs. 👇
🩺 Demo: A clinical data pipeline powered by Ling-3.0-flash, running entirely on one DGX Spark. 🔒 Local de-identification keeps patient data private 🤖 Extracts vitals, lab results and ultrasound findings, then assigns study groups 🚩 Flags ambiguous cases for human review
Download the weights: 🤗 Hugging Face: FP4: https://t.co/hRGMeO7sUo INT4: https://t.co/UBzvHyVVnm 🤖 ModelScope: FP4: https://t.co/bfhiJSlGQM INT4: https://t.co/LnJICZWmnV
Four-bit weights put a 124B on a single DGX Spark at roughly 80 per second decode, making fully local coding agents and offline batch practical on one box.
Checking sign-in…
Loading comments…