Meet Ling-3.1-flash: ~560B total params, ~25B active/token, up to 1M-token context.
We plan to open-source the model soon.
Across work, coding & healthcare: 1,673 Elo on GDPVal-AA v2.1, 75.16 on FrontierSWE, and 65.35 on HealthBench Professional.
mixture-of-experts — A model built from many specialist sub-networks where only a few activate per token, giving big-model capability at small-model running cost.
token — The chunk of text a model reads and writes in — roughly three-quarters of a word — and the unit AI usage is billed in.
context window — The maximum amount of text a model can consider at once — its working memory for the current conversation or task.
open weights — A model whose trained parameters are published for anyone to download and run — unlike API-only models you can access but never possess.
Why it matters
Ant Group's Ling-3.1-flash is a 1M-context windowThe maximum amount of text a model can consider at once — its working memory for the current conversation or task.Full definition →mixture-of-expertsA model built from many specialist sub-networks where only a few activate per token, giving big-model capability at small-model running cost.Full definition → with agentic coding results near Claude Fable 5 on long tasks like a compiler build and a Pyright speedup. open weightsA model whose trained parameters are published for anyone to download and run — unlike API-only models you can access but never possess.Full definition → are promised, so watch for a release.