Ling-3.0-tiny is now available as an open-weight model in BF16, FP8 and INT4. On Artificial Analysis, it scores 25 on the Intelligence Index and 16 on the Agentic Index, with 772 Elo on GDPval-AA v2 and 20.80 on τ³-Banking—built for real task execution. 🧵


Across broader evaluations, Ling-3.0-tiny reaches 71.03 on IMO-AnswerBench, 73.40 on GPQA Diamond, 83.15 on Multi-IF and a 69.54 non-hallucination rate. It delivers balanced coverage across reasoning, coding agents, instruction following and long-context tasks.

Efficiency is architectural: a 3:1 KDA–MLA hybrid attention stack and 128 sparse MoE experts, with 8 routed and 1 shared expert active per token. The result is high model capacity with lower per-token compute and a more accessible deployment footprint.

Already running on real hardware: • DGX Spark FP8: ~100–105 tok/s for one request; ~161 tok/s aggregate for two • MacBook: local BF16/FP8, with FP8 ~30% faster in sustained generation • Mac mini: an always-on node for private knowledge, automation and local APIs
A 7.9B model with 1.3B active parameters and published agentic scores runs locally on a MacBook or DGX Spark, giving a private, zero-API-cost option for tool-use and retrieval loops.
Checking sign-in…
Loading comments…