Today we release Pipette, a model evaluation suite for on-device intelligence, in partnership with @ArtificialAnlys Most benchmarking platforms are optimized to measure core capabilities and speed profile of foundation models served in the cloud. Pipette gives the field a common, reproducible way to measure the quality, speed, latency, and memory use of AI models on devices such as phones, laptops, PCs, AI boxes, and embedded hardware. > Pipette is open source > In Pipette, models get compated as model + quantization + runtime + device from one interface. > It comes with a warehouse of verified benchmark results, currently with 10k+ results across 35 model classes, 7 quants, llama.cpp runtimes, and 4 devices. > Pipette is a dynamic platform, allowing new contributions from day one, adding new devices, runtimes, model families, and quantization levels. 🧵

What is being released today: 1. An interactive dashboard for exploring and comparing on-device model performance: https://t.co/1iLOhrbPfA 2. A public dataset containing more than 10k results, including 1000+ tested performance configurations across models, quantization levels, runtimes, and devices. Inspect the current breakdown: https://t.co/VquWJOeje4 3. Benchmark clients for macOS, Windows, iOS, and Android that run versioned benchmark definitions on target devices: https://t.co/VbciJQTGsv 4. A companion Artificial Analysis dashboard presenting Pipette results and the aggregate benchmark scores: https://t.co/83tOaaX7z0 5. Apache 2.0-licensed open-source infrastructure for benchmark management, execution, submissions, and evaluation scoring: https://t.co/kk8jnHRtAf https://t.co/VbciJQTGsv https://t.co/9He1rFbafZ 6. Native iOS and Android apps for running benchmarks directly on-device. iOS App: https://t.co/XcxCabRNZO Android app: https://t.co/jIFlM43285 2/
Get started: The initial public release covers about 35 model classes from several providers, with 7 llama.cpp quantization levels. Laptop and desktop runs cover all four levels, while current phone runs cover all quant levels.. The release includes context lengths from 256 to 8,192 tokens, where device memory allows. Current devices include a MacBook Pro with M5 Max, iPhone 17 Pro, and Samsung Galaxy S26 Ultra, AMD Ryzen AI Max+ 395 with Radeon 8060S (coming live soon). More models, runtimes, and devices are on the way. > If you are choosing a model, inspect the runtime, quantization, and device you plan to ship. > If that configuration is missing, run the Pipette clients on your hardware and help test the third-party submission workflow. > Every accepted submission expands the configurations represented in the dataset. > We encourage model providers, runtime developers, hardware teams, and application developers to test the configurations they care about. > If there is something you want measured, open an issue and tell us. 3/
Read the full article: https://t.co/NH4MK3I88c Explore results in Artificial Analysis: https://t.co/83tOaaX7z0 > Explore the Pipette dashboard https://t.co/1iLOhrbPfA > Read the methodology https://t.co/lnl0zCSLJF > Request a configuration or report an issue https://t.co/3E9a7cwGra 4/4

Before shipping a local model you can check the exact , runtime and device you plan to target, or run the clients on your own hardware and submit the configuration when it is missing from the dataset.
Checking sign-in…
Loading comments…