
Ant Ling
29 Intel
Ant Ling AGI is the account for Ant Group's AGI initiative, publishing the Ling and Ring families of open-source large language models built on a mixture-of-experts architecture. The Ling series are general-purpose models and the Ring series are reasoning-focused, with releases such as Ring-1T reaching trillion-parameter scale and performing strongly on IMO-style reasoning benchmarks. The group also publishes the multimodal Ming model family for image, text, audio, and video.
Is this you? Sign in with X to claim this profile.
Intel
Ant Group open-sourced Ling-3.0-flash-Fin, a 124B-parameter MoE with 5.1B active parameters and a 256K context window, released on Hugging Face, along with FinFIRST, a financial search agent benchmark of 123 expert-authored tasks, 701 atomic criteria and 12,300 rubric points built with CICC.
Ant Group ships FP8, FP4 and INT4 quantized checkpoints of Ling-3.0-flash-Fin on Hugging Face and ModelScope, letting teams pick a precision tier that fits their memory and latency budget for financial-workflow deployments.
Ant Ling post-trains Ling-3.0-tiny with GSPO on a single DGX Spark using the AReno framework. Over 400 steps on a tic-tac-toe environment, mean reward moved from about -0.5 to 0.4 and responses shortened to roughly 850 tokens, with more stable tool calls. Tutorial and code are published.
Ant Ling open-sourced the Ming-Image-0.1-Design family (6B parameters) plus two agent skills for UI design and image-to-editable-PPT conversion; the base model ranks #1 among open-weight models on Artificial Analysis's UI/UX Design leaderboard.
Ant's Ling team launches Ring-2.6-1T, a trillion-parameter reasoning model with a selectable thinking-effort dial. It posts strong agentic execution and tool-orchestration scores, targets multi-step production workflows, and is free on OpenRouter for a week with open weights promised.
Ant Ling's successor to IcePop swaps the fixed-ratio token mask for an adaptive binary-KL trust region matched to each token's inherent noise, which the team credits for stable trillion-scale long-horizon agentic RL and a pure-RL SWE-bench Verified score above 76.
Ant Group's Ling-3.1-flash is a ~560B MoE with ~25B active parameters and a 1M-token context. Self-reported results include multi-hour coding runs such as a Lua-to-x86 compiler. Open-sourcing is planned.
Ant Ling published six base checkpoints for Ling-3.0-tiny (7.9B total, 1.3B active) and Ling-3.0-flash (124B total, 5.1B active), spanning pretrained, mid-trained and WSM-merged stages with no post-training, alongside the paper on replacing learning-rate decay with weighted checkpoint merging.
Ant Group's hybrid-reasoning MoE stacks KDA and MLA layers 5:1 with 1/64 expert activation, supports 256K context natively and scales to 1M, and is claimed to match or beat the team's 1T flagship on most reported benchmarks. Free on OpenRouter through August 3, 2026.
Ant Group releases Ling-3.0-flash-Sante, a health- and medicine-focused MoE model built on Ling-3.0-flash, reporting competitive results on MedXpertQA-Text, DiagnosisArena-MCQ, HealthBench Professional and BrowseComp against flagship models.
Ant Group's small hybrid-reasoning MoE arrives on OpenRouter and Vercel's AI Gateway with a free window to August 13, alongside demos covering local Obsidian retrieval, offline translation, and native tool use for mobile and browser control. Open weights promised to follow.
Ant Ling open-sourced DSpark, a speculative decoding draft model for Ling-3.0-flash reaching 1,120 tok/s and 0.78 ms mean TPOT at batch 1 on four Blackwell GPUs. The thread covers the draft architecture, acceptance-aware training, and removing a FlashInfer sync that left steps 47% idle.
Ling-3.0-flash-VL moved from free to paid access on OpenRouter as of September 23, with discounted introductory pricing of $0.021 input / $0.063 output / $0.0042 cache-hit per million tokens.
Ant Group's 7.9B MoE with 1.3B active parameters is downloadable in BF16, FP8 and INT4, with published Artificial Analysis, GPQA and IMO-AnswerBench scores, a 3:1 KDA/MLA attention stack, 128 sparse experts, and measured local throughput on DGX Spark and Apple silicon.
Ling-3.0-flash-VL is now free on OpenRouter for two weeks, and Ant Group has open-sourced FP4 and INT4 quantized versions for self-hosted multimodal agent deployments.
Working with the SGLang team, Ant Ling built a Fused MoE V2 Pallas kernel for TPU v7x that overlaps token routing and HBM weight prefetch with compute, reporting 53% lower MoE prefill latency and up to 1.77x decode throughput on a 16-chip slice against a comparable H200 cluster.
INT4 and MXFP4 builds of the 124B MoE run end to end on one DGX Spark through an adapted SGLang path, measured at about 80 tokens/s decode, 2,500 to 3,500 tokens/s prefill, and three or four concurrent users. W4A16 is the stable FP4 default, W4A8 the throughput option.
The 124B MoE with roughly 5.1B active parameters per token is now downloadable from Hugging Face and ModelScope in BF16 and an FP8 quantization, letting teams run their own evaluations and self-host it inside agentic toolchains rather than calling a hosted endpoint.
Ant introduces Ling-3.0-flash-Fin, a 124B/5.1B-active finance variant of Ling-3.0-flash built with financial institutions. Demos run 20 to 55 tool calls across filings and multi-sheet workbooks with traceable evidence, and its general index score rose from 38 to 41. Weights open next week.
Artificial Analysis benchmarked Ling-3.0-tiny on phone-ready quantized builds, scoring 59 at 16K context with 5.7 seconds on an iPhone 17 Pro and top rank at 64K with 66. A rare case of intelligence and real-device latency measured on the same build.
A persistent per-GPU daemon holds already-sharded, already-quantized weights in memory and serves CUDA IPC handles to restarting SGLang engines. Ling-2.6-1T FP8 startup falls from 8.8 minutes to about 0.53, with config validation guarding correctness and a disk fallback on mismatch.
Ant Ling reports that E2M1's non-uniform 4-bit grid causes a shrinkage bias under RTNE rounding that accumulates across GEMMs and layers. A uniform E1M2 or INT4 style grid removes it and stays closer to BF16 loss across 1.5B dense, 7.9B MoE and 124B MoE pretraining runs.
SwiGLU grows quadratically for large inputs, inflating activations and outliers. PowLU decays that growth toward linear with one hyperparameter, matches SwiGLU scaling-law exponents from 26M to 368M activated parameters, and avoided the FP8 loss spikes both SwiGLU variants hit near step 77k.
The report details a 7:1 Lightning Attention to MLA hybrid that makes 256K context affordable, plus KPop, an adaptive binary-KL replacement for uniform KL constraints in MoE RL that lifted SWE-bench Verified from 70.8% to 76.28%. Base weights for the flash and 1T models ship with it.
Ant Group trained a trillion-parameter model straight from base with RL and no extra human annotation, reporting that hand-crafted format rewards become redundant at that scale. The four-stage elicit, distill, refine and adaptive pipeline produces shorter correct chains and unusually strong distilla
Ant Ling confirmed Ling-3.0-flash-fin is actually a 124B-parameter MoE model with 5.1B active parameters despite the "flash" name, and noted an fp4 quantization is available for local deployment.
Ling-2.6-1T is now callable on OpenRouter. The trillion-parameter instruct model skips verbose reasoning traces to cut token cost by roughly 75% while claiming to hold its ground on AIME26 and SWE-bench Verified, aimed at coding and large agent workflows.
Ant Group details FinFIRST, a finance-agent benchmark built with 50+ finance professionals using atomic rubrics to grade both answers and evidence quality, benchmarking 15 model configurations under a unified search/browse/Python setup.
The trillion-parameter thinking model ships with high and xhigh reasoning modes, separating production agent loops from deep reasoning, trained with asynchronous RL via IcePop. Reported results include 74.00 SWE-bench Verified, 95.83 AIME 26, 88.27 GPQA Diamond and 66.18 ARC-AGI-V2 pass@2.