Vibeleaderboard
← All Intel
Intel / post

Ling-3.0-flash: a 124B MoE with 5.1B active parameters for agents

Source
Ant Ling
Date
Ant Ling@AntLingAGI
Thread · 12 parts

Today, we’re releasing Ling-3.0-flash—a hybrid-reasoning MoE model built for production-scale agents. 124B parameters. Just 5.1B active per token. With 1/8 of the total and 1/12 of the active parameters, it matches or beats our 1T flagship model on most benchmarks shown.

Ling-3.0 starts with native hybrid-linear attention: KDA and MLA layers stacked 5:1. KDA gives fine-grained control over long-range memory, while 1/64 expert activation makes MoE compute more efficient. It supports 256K context natively and can scale to 1M.

Ling-3.0-flash is now live on OpenRouter—and free to use through August 3, 2026. Try it in your coding, search, research, and tool-use workflows. Then show us what you build: https://t.co/fOgz3YCqVc

Demo 1 — From one prompt to a 3D world. Using Blender MCP, Ling-3.0-flash wrote Python, built a city with elevated roads, skyscrapers, and materials, set the camera path, and rendered an aerial video—showing spatial reasoning and long-horizon tool use.

Read the full thread on X

Context

Ant Ling releases Ling-3.0-flash, a 124-billion-parameter model that activates just 5.1 billion parameters per token, roughly one-eighth the total size and one-twelfth the active parameters of the team's trillion-parameter flagship. The company says the smaller model matches or beats that flagship on most of the benchmarks it shows, though the specific benchmarks and scores behind that comparison are not listed in the thread. Architecturally, Ant Ling describes stacking KDA and MLA layers five to one, with only 1 in 64 experts activated per token, and says this supports 256,000 tokens of natively with the ability to scale to 1 million.

The thread includes seven demonstrations the company says show the model building a 3D city in Blender through tool calls, coordinating five separate roles to research and write a paper, formatting a Word proposal and checking Excel financials through an Office integration, and completing a calendar-scheduling workflow end to end. These were shown as video walkthroughs and are described here only as the company presented them in writing, not independently assessed. Ling-3.0-flash launched free on OpenRouter through August 3, 2026, with the company saying an open-source release would follow.

Terms in this piece · Glossary
  • mixture-of-experts — A model built from many specialist sub-networks where only a few activate per token, giving big-model capability at small-model running cost.
  • attention — The mechanism that lets a model weigh which earlier words matter for the word it's currently processing — the core operation of a transformer.
  • context window — The maximum amount of text a model can consider at once — its working memory for the current conversation or task.
  • AI agent — An AI system that doesn't just answer once but works toward a goal in a loop — taking actions, reading the results, and deciding what to do next.
More from Ant Ling
Recommended reads
Comments

Checking sign-in…

Loading comments…