Vibeleaderboard
← All Intel
Intel / post

Ling 3.1 Flash doubles its Intelligence Index score to 41

Source
x.com
Date
A
ArtificialAnlys@ArtificialAnlys

Ling 3.1 Flash makes large gains in intelligence over its predecessor, scoring 41 on the Artificial Analysis Intelligence Index with particular improvement in agentic capabilities @AntLingAGI has released Ling 3.1 Flash, a reasoning model with 560B total parameters, 25B active parameters, and a 1M token context window. It scores 41 on the Artificial Analysis Intelligence Index v4.3, up from 20 from Ling 3.0 Flash, which launched in August. Ling 3.1 Flash is expected to be open weights, with Ant Group releasing the weights soon. Key results: ➤ Ling 3.1 Flash improves agentic capabilities over its predecessor. With a GDPval-AA v2 Elo of 1,622 and an AA-Briefcase Elo of 1,400, the model rivals peer models such as GLM-5.3-Flash, Gemini 3.8 Flash (high) and DeepSeek V4.1 Flash (Max). Ling 3.1 Flash also shows notable gains on AutomationBench-AA (62%) and Terminal-Bench v4.0 (33%). ➤ Ling 3.1 Flash scores +2 on AA-Omniscience, a 22 point improvements from Ling 3.0 Flash (-18). Compared to its predecessor, its accuracy rate rose from 18% to 29% while the hallucination rate fell from 44% to 38% at a similar attempt rate (56% to 58%), demonstrating the gain comes primarily from knowing…

Read the full post on X
Why it matters

Ling 3.1 Flash (560B total, 25B active, 1M ) jumps from 20 to 41 on the Intelligence Index, with gains on agentic and terminal tasks. Weights are expected soon, which makes it a candidate open model for agent workloads.

Key takeaways · AI-distilled
  • Artificial Analysis reports Ling 3.1 Flash scores +2 on AA-Omniscience, up from -18: accuracy rose from 18% to 29% and fell from 44% to 38% at a similar attempt rate, so the gain comes from knowing more, not abstaining more.
  • On agentic tests it posts a GDPval-AA v2 Elo of 1,622 and an AA-Briefcase Elo of 1,400, plus 62% on AutomationBench-AA and 33% on Terminal-Bench v4.0, which Artificial Analysis says rivals GLM-5.3-Flash and Gemini 3.8 Flash.
  • It is far larger than Ling 3.0 Flash (124B total, 5.1B active) and pricier at $0.30/$0.90 per 1M input/output tokens, yet used 16% fewer output tokens (218M vs 261M) to run the Intelligence Index.
  • At $0.99 per task it costs well above GLM-5.3-Flash ($0.42) and DeepSeek V4.1 Flash ($0.32) but below Gemini 3.8 Flash (High) at $1.24. It is available through Novita AI, with weights coming soon.
Terms in this piece · Glossary
  • reasoning model — A model trained to think — generating extended internal reasoning before answering — trading time and tokens for accuracy on hard problems.
  • context window — The maximum amount of text a model can consider at once — its working memory for the current conversation or task.
  • hallucination — When a model states something false with full confidence — inventing facts, citations, or APIs that don't exist.
  • open weights — A model whose trained parameters are published for anyone to download and run — unlike API-only models you can access but never possess.
More from ArtificialAnlys
Recommended reads
Comments

Checking sign-in…

Loading comments…