Today, we’re releasing Ling-3.0-flash—a hybrid-reasoning MoE model built for production-scale agents. 124B parameters. Just 5.1B active per token. With 1/8 of the total and 1/12 of the active parameters, it matches or beats our 1T flagship model on most benchmarks shown.

Ling-3.0 starts with native hybrid-linear attention: KDA and MLA layers stacked 5:1. KDA gives fine-grained control over long-range memory, while 1/64 expert activation makes MoE compute more efficient. It supports 256K context natively and can scale to 1M.

Ling-3.0-flash is now live on OpenRouter—and free to use through August 3, 2026. Try it in your coding, search, research, and tool-use workflows. Then show us what you build: https://t.co/fOgz3YCqVc
Demo 1 — From one prompt to a 3D world. Using Blender MCP, Ling-3.0-flash wrote Python, built a city with elevated roads, skyscrapers, and materials, set the camera path, and rendered an aerial video—showing spatial reasoning and long-horizon tool use.
A 124B that activates only 5.1B parameters per and holds 256K natively targets long-horizon at close to small-model serving cost, worth benchmarking against larger hosted models.
Checking sign-in…
Loading comments…