Ling-3.0-tiny Measured on the Phone-Class Intelligence and Speed Frontier
- Source
- Ant Ling
- Date
According to Artificial Analysis, Ling-3.0-tiny sits on the mobile intelligence–speed Pareto frontier: 59 at 16K and 5.7s on iPhone 17 Pro. It also ranks first in the 64K intelligence evaluation with a score of 66. Bringing stronger intelligence to smaller devices.


Thanks to @ArtificialAnlys for evaluating intelligence and real-device inference on the same phone-ready quantized builds. Explore the full benchmark and methodology: https://t.co/vHGpiekGRk
Context
Ant Ling reports that Artificial Analysis, an independent benchmarking group, placed its small Ling-3.0-tiny model on the Pareto frontier for phone-class , meaning no other phone-ready model it tested beat Ling-3.0-tiny on both intelligence and speed at the same time. The specific figures cited are a score of 59 at a 16K-token evaluation with a 5.7-second response on an iPhone 17 Pro, and a first-place score of 66 on Artificial Analysis's 64K-token intelligence evaluation among the models it compared.
Artificial Analysis had separately benchmarked the model two weeks earlier, reporting an Intelligence Index score of 25 for the 7.9-billion-parameter, 1.3-billion-active model at 154 tokens per second, well above the median for models its size, but noting it uses roughly four times the median to reach that score. Read together, the two results describe a model that trades token efficiency for accuracy relative to peers its size, a tradeoff a mobile deployment may or may not want depending on whether latency or per-token cost matters more.
- inference — Running a trained model to get answers — the phase where AI is actually used, as opposed to trained.
- quantization — Shrinking a model by storing its numbers less precisely — like rounding — so it runs faster and fits on smaller hardware, at a small quality cost.
- token budget — A cap on how many tokens a task, session, or agent run may consume — the practical control on both cost and how long an agent will grind.
- context window — The maximum amount of text a model can consider at once — its working memory for the current conversation or task.
Checking sign-in…
Loading comments…






