Today, we’re releasing the open weights for Ling-3.0-flash. 🎉 Official BF16 and FP8-quantized versions are now available, so you can choose the option that best fits your hardware, performance requirements, and deployment needs.
Get the Ling-3.0-flash weights here: 🤗 Hugging Face Official https://t.co/rCkfU23Li6 FP8 https://t.co/I6mZ1WhNQN 🤖 ModelScope Official https://t.co/hN7GJygYtG FP8 https://t.co/h7e4OGeWo6
With the weights in your hands, you’re free to take Ling-3.0-flash further. Run your own evaluations. Deploy it in your own environment. Adapt it to your use cases. Build it into your products, development toolchains, and agentic workflows. Download it, make it yours, and show us what you build. 🚀
A huge thank-you to the @sgl_project and @vllm_project communities for their continued, high-quality support! Your contributions have helped make Ling-3.0-flash easier and more efficient to deploy across diverse hardware and production environments. 🚀
Self-hosting a 124B with 5.1B active parameters becomes possible: BF16 and FP8 weights on Hugging Face and ModelScope let you run private evaluations and deployments instead of depending on a provider endpoint.
Checking sign-in…
Loading comments…