Ant Group's finance model matches MiniMax-M2.7 with half the parameters
- Source
- ArtificialAnlys
- Date
Ling-3.0-flash-Fin, Ant Group’s new finance-focused open weights model, scores 23 on the Artificial Analysis Intelligence Index and 24 on the Finance & Accounting Index, and is on the Intelligence vs. Active Parameter Pareto Frontier @AntLingAGI has released Ling-3.0-flash-Fin, a finance-focused model built on Ling-3.0-flash. Ant Group announced that it developed the model with financial institutions and industry experts to support financial research, including checking sources, building valuation spreadsheets and writing reports. This text-only model comes after their release of their image and video input-capable model Ling-3.0-flash-VL, which scored 25 on the Intelligence Index. Key results: ➤ Ling-3.0-flash-Fin matches MiniMax-M2.7’s Intelligence Index score with roughly half the active parameters. Both score 23, while Flash-Fin activates 5.1B parameters per token compared with MiniMax-M2.7’s 10B. ➤ Ling-3.0-flash-Fin matches Ling-3.0-flash-VL at 24 on the Artificial Analysis Finance & Accounting Index. Fin has higher business knowledge accuracy than VL (17% vs. 11%), but also higher business knowledge hallucination (33% vs. 19%) ➤ Ling-3.0-flash-Fin scores slightly below Ling-3.0-flash-VL on professional knowledge work. It scores 1171 Elo on GDPval-AA v2 and 967 on AA-Briefcase, compared with 1225 and 986 respectively for Ling-3.0-flash-VL. Both benchmarks test agents on professional tasks such as producing documents and spreadsheets. ➤ Difficult agentic tasks remain a challenge for Ling-3.0-flash-Fin. It scores 7% on AutomationBench-AA, which tests workflows across business apps while respecting guardrails, compared with 16% for the flash-VL model. Both models score 0% on Terminal-Bench v4.0, which tests difficult terminal-use tasks. ➤ Ling-3.0-flash-Fin uses more output tokens than the flash-VL model and MiniMax-M2.7. It averages ~67k output tokens per Intelligence Index task, about 34% more than VL (~50k) and 3.2x MiniMax-M2.7 (~21k). Additional model details: ➤ Type: Open weights reasoning model. ➤ Size: 124B total parameters, 5.1B active per token (MoE). ➤ Context window: 256K tokens. ➤ Modalities: Text input and output. ➤ API availability: Available through @OpenRouter, including a rate-limited free endpoint. ➤ License: MIT.

Ling-3.0-flash-Fin sits on the Intelligence Index vs. active parameters Pareto frontier, scoring 23 with 5.1B active parameters per token vs. a score of 25 or Ling-3.0-flash-VL which uses 5.5B active per token. Both have 124B total parameters, while Qwen3.8 27B (xhigh) scores 34 with 27B total parameters, placing both Ling models below the total-parameter frontier.

Ling-3.0-flash-Fin scores 24 on the Artificial Analysis Finance & Accounting Index, which combines business knowledge, reasoning, agentic work, long-context analysis and non-hallucination. This is the same score as Ling-3.0-flash-VL both. Fin has higher business knowledge accuracy (17% vs. 11%), but also worse business knowledge hallucination (non-hallucination rate of 67% vs. 81%).

Ling-3.0-flash-Fin scores 1171 Elo on GDPval-AA v2, which tests agents on professional knowledge work. This is ~50 points behind Ling-3.0-flash-VL which scored 1225, and is above MiniMax-M2.7, which scored 1087.

Context
Ant Group's Ling-3.0-flash-Fin, a finance-focused open-weights model built on Ling-3.0-flash, scores 23 on the Artificial Analysis Intelligence Index, the same as MiniMax-M2.7, according to Artificial Analysis, a third-party benchmark tracker. Fin does it with 5.1 billion active parameters per token against MiniMax-M2.7's 10 billion. It is a model, so only part of its 124 billion total parameters runs for each token.
The tie is narrower than it looks. Against Ant's vision-capable Ling-3.0-flash-VL, Fin scored lower on professional knowledge work (1171 versus 1225 Elo on GDPval-AA v2) and on AutomationBench-AA, which tests business-app workflows under (7% versus 16%). Both scored 0% on Terminal-Bench v4.0. Fin also showed higher business-knowledge (33% versus 19%) and averaged about 67,000 output tokens per Intelligence Index task, roughly 3.2 times MiniMax-M2.7's 21,000. An index score alone does not show those differences.
- mixture-of-experts — A model built from many specialist sub-networks where only a few activate per token, giving big-model capability at small-model running cost.
- guardrails — The checks around a model that block bad inputs and outputs — filters, validators, and permission rules the model itself can't override.
- hallucination — When a model states something false with full confidence — inventing facts, citations, or APIs that don't exist.
- open weights — A model whose trained parameters are published for anyone to download and run — unlike API-only models you can access but never possess.
Checking sign-in…
Loading comments…






