Which model should you run on the iPhone 17 Pro? Announcing our new intelligence and inference testing for small models on mobile devices: independent measurement of how capable small models are in typical on-device tasks, and how they perform on popular phones - in partnership with @liquidai We have partnered with @liquidai to deliver mobile device inference benchmarking, covering a range of models in 4-bit or lower precision on the iPhone 17 Pro and Galaxy S26 Ultra We’re publishing our phone-scale intelligence evaluation results in combination with @liquidai's inference performance benchmarks to give users and developers a holistic view of how small models are performing on phones Inference benchmarking is conducted in a controlled environment using Liquid AI’s inference benchmarking software, which Artificial Analysis has examined and is open-sourced on Liquid AI’s GitHub. Results cover end-to-end generation time, output speed, peak memory usage and other metrics. The inference benchmarking app, ‘Pipette’, is available to download for free on iOS and Android, allowing users to test a variety of models on their own devices We are ranking phone-scale model intelligence based on each model’s average score in five evaluations chosen for the task-based work these models do in practice: BFCL, IFBench, AA-Omniscience, GPQA Diamond and MATH-500. These evaluations are run by Artificial Analysis using our independent methodology. By default, we limit models to 16K context on each of these evaluations, representing the lack of memory space for significant KV cache on mobile devices. This leads to some intelligent but verbose models dipping in relative score - they were not designed for the constraints that phone memory imposes on token use We are defining our portable device category as including models that fit within 8 GB of memory after quantization, including KV cache, at 8K context We expect both the intelligence and inference benchmarks to evolve over time, as new models, devices, inference frameworks and quantization techniques are released. Our pages will remain up to date with these new additions Initial results: ➤ Nanbeige4.2-3B and LFM2.5-2.6B share the top average evaluation score at 63 (with a 16K context limit), ahead of Ornith-1.0-9B at 62 and Qwen3.5 9B (Reasoning) at 61. LFM2.5-2.6B achieves its score more efficiently: on an iPhone 17 Pro it answers a standard 1,024-token prompt in 8.0s using 2.3 GB of memory, against 21.4s and 4.0 GB for Nanbeige4.2-3B and 25+ seconds and 6.9 GB for the two 9B models ➤ The 16K context limit shapes the leaderboard: as an example, Qwen3.5 9B (Reasoning) spends 74.5M output tokens across one pass of the benchmark set, hitting the 16K limit on 29% of its generations and landing in fourth place overall. With the limit raised to 64K, Ling 3.0 Tiny takes first place with a score of 66, ahead of Nanbeige4.2-3B (65), Qwen3.5 9B (Reasoning, 64), and LFM2.5 2.6B (64). But a 64K window does not fit in mobile phone memory, and at 55 output tokens/s on an iPhone, generating 64K tokens could mean a 20+ minute wait and a lot of battery use. This is why our primary results are capped at 16K, but we're also publishing a set of results capped at 64K, and another capped at one minute of generation time ➤ The speed-intelligence Pareto frontier is short: six models are unbeaten on both intelligence and speed on an iPhone 17 Pro: LFM2.5-230M (27 at 0.9s), MiniCPM5-1B (45 at 2.9s), LFM2.5-8B-A1B (58 at 5.7s), Ling 3.0 Tiny (59 at 5.7s), LFM2.5-2.6B (63 at 8.0s) and Nanbeige4.2-3B (63 at 21.4s). LFM2.5-8B-A1B and Ling 3.0 Tiny are mixture-of-experts models that activate ~1B parameters per token, which is how they answer in under 6s with 8B-class weights ➤ Leading models have opposite strengths: Nanbeige4.2-3B is the most balanced (76% on BFCL, 96% on MATH-500, 67% on GPQA Diamond); Qwen3.5 9B (Non-reasoning) is the strongest tool caller (77% on BFCL) and scientific reasoner (79% on GPQA Diamond); LFM2.5-2.6B follows instructions best of any model measured on the iPhone (59% on IFBench), clears 90% on MATH-500 and hallucinates far less on AA-Omniscience (79% non-hallucination, against 33% for Nanbeige4.2-3B and 24% for Qwen3.5 9B (Reasoning)) More details below in thread ⬇️

At the standard 16K context limit, Nanbeige4.2-3B (Reasoning) and LFM2.5-2.6B (Reasoning) tie for the top average evaluation score at 63, ahead of Ornith-1.0-9B (Reasoning, 62), Qwen3.5 9B (Reasoning, 61), Ornith-1.5-9B (Reasoning, 61), Gemma 4 E4B (Reasoning, 60) and Qwen3.5 9B (Non-reasoning, 60). Ornith-1.5-9B slips behind its older 1.0 sibling purely on a weaker instruction following performance in IFBench When the context limit is raised to 64K (represented by dots in the image), Ling 3.0 Tiny takes the top spot at 66, followed by Nanbeige4.2-3B at 65, and both Qwen3.5 9B (Reasoning) at 64 and LFM2.5-2.6B at 64

In the individual evaluations: Qwen3.5 9B (Non-reasoning) takes BFCL at 77% and GPQA Diamond at 79%, Falcon-H1R-7B takes MATH-500 at 97%, LFM2.5-2.6B takes IFBench at 59% and hallucinates far less than the other leading models on AA-Omniscience (79% non-hallucination), Ornith-1.0-9B recalls the most facts (15% accuracy), and G9v3-3B has the highest non-hallucination rate of any model that attempts answers (89%). Nanbeige4.2-3B wins none of the evaluations outright but is top-five on BFCL, GPQA Diamond and MATH-500 - which is how it ties for first overall

A score of ~60 can cost 5M output tokens or 75M. Qwen3.5 9B (Reasoning) spends 74.5M output tokens across one pass of the benchmark set and runs out of its 16K window on 29% of generations; Gemma 4 E4B (Reasoning) reaches a similar overall score on 5.2M tokens and never hits the limit. On a phone, token use leads to time, energy use and heat that are less tolerable than on a laptop, desktop or server

Choosing an on-device model is a three-way trade: two models can tie on score while one burns 14x the output and 6.9GB of RAM. These numbers make that trade explicit per handset.
Checking sign-in…
Loading comments…