Vibeleaderboard
← All Intel
Intel / post

Ling-3.0-tiny ships open weights in BF16, FP8 and INT4

Source
Ant Ling
Date
Ant Ling@AntLingAGI
Thread · 6 parts

Ling-3.0-tiny is now available as an open-weight model in BF16, FP8 and INT4. On Artificial Analysis, it scores 25 on the Intelligence Index and 16 on the Agentic Index, with 772 Elo on GDPval-AA v2 and 20.80 on τ³-Banking—built for real task execution. 🧵

Across broader evaluations, Ling-3.0-tiny reaches 71.03 on IMO-AnswerBench, 73.40 on GPQA Diamond, 83.15 on Multi-IF and a 69.54 non-hallucination rate. It delivers balanced coverage across reasoning, coding agents, instruction following and long-context tasks.

Efficiency is architectural: a 3:1 KDA–MLA hybrid attention stack and 128 sparse MoE experts, with 8 routed and 1 shared expert active per token. The result is high model capacity with lower per-token compute and a more accessible deployment footprint.

Already running on real hardware: • DGX Spark FP8: ~100–105 tok/s for one request; ~161 tok/s aggregate for two • MacBook: local BF16/FP8, with FP8 ~30% faster in sustained generation • Mac mini: an always-on node for private knowledge, automation and local APIs

Read the full thread on X

Context

Ant Group's Ling team released Ling-3.0-tiny as in three numeric precisions: BF16 (full precision), FP8, and INT4 (heavily compressed for smaller, faster deployment). The model has 7.9 billion parameters in total, but it is a design, meaning only a fraction of those parameters, about 1.3 billion, actually run for any given token, as Ant also states in a companion post about the same model. That keeps compute cost closer to a much smaller dense model while retaining more of a larger model's stored capability.

In its release thread, Ant says the evaluator Artificial Analysis scored the model at 25 on its Intelligence Index and 16 on its Agentic Index, and cites a 772 Elo rating on GDPval-AA v2 and 20.80 on the banking- benchmark tau3-Banking. Those are third-party benchmark labels as Ant describes them in its own thread, alongside further scores of 71.03 on IMO-AnswerBench and 73.40 on GPQA Diamond.

On real hardware, Ant reports roughly 100 to 105 tokens per second for a single request on an NVIDIA DGX Spark running the FP8 version, rising to about 161 tokens per second when two requests run together, and it also ran the BF16 and FP8 versions locally on a MacBook, with FP8 about 30% faster than BF16 in sustained generation. In a local demo pairing the model with a clickable-word reference tool, Ant reports the first generated token typically arrived in under 100 milliseconds, illustrating what that speed feels like offline and without an API call.

Terms in this piece · Glossary
  • open weights — A model whose trained parameters are published for anyone to download and run — unlike API-only models you can access but never possess.
  • mixture-of-experts — A model built from many specialist sub-networks where only a few activate per token, giving big-model capability at small-model running cost.
  • AI agent — An AI system that doesn't just answer once but works toward a goal in a loop — taking actions, reading the results, and deciding what to do next.
  • attention — The mechanism that lets a model weigh which earlier words matter for the word it's currently processing — the core operation of a transformer.
More from Ant Ling
Recommended reads
Comments

Checking sign-in…

Loading comments…