Skip to main content

Liquid AI Open-Sources Pipette Benchmarking Suite

Liquid AI partners with Artificial Analysis to launch Pipette, an Apache 2.0 open-source suite for measuring on-device AI models.

AI-written
Inewgen
26 Aug 2026Source: MarkTechPost3 min read (0 views)Last updated 29 Aug 2026
Share
Liquid AI Open-Sources Pipette Benchmarking Suite

Stock photo for illustration only, not from the actual event

Font size
  • Liquid AI releases Pipette, an open-source benchmarking suite under the Apache 2.0 license
  • Developed in partnership with Artificial Analysis to review and verify methodology
  • Launch dataset covers five metrics across more than 1,000 on-device configurations
  • Supports over 30 models, various quantization formats, and context lengths up to 8,192 tokens

Liquid AI has made waves in the artificial intelligence community by officially releasing Pipette, an open-source benchmarking suite designed to evaluate on-device AI model performance. Distributed under the Apache 2.0 license, the infrastructure includes pipette-mgmt, pipette-clients, and pipette-scores, alongside a public results dataset, a hosted dashboard, and native iOS and Android benchmark apps with no waitlist required, though community-submitted results remain in beta.

The release was developed in partnership with Artificial Analysis, an independent validator that reviewed and verified the benchmarking methodology. The core premise is both narrow and practical: on-device behavior is a property of the deployed system rather than the model in isolation. The initial launch dataset covers five on-device performance metrics across more than 1,000 configurations combining models, quantization formats, runtimes, devices, and context lengths.

1,000+Benchmark Configurations
30+Supported Models

The suite supports over 30 models, multiple quantization formats, llama.cpp builds for macOS, iOS, Windows, and Android, and context lengths ranging from 256 to 8,192 tokens. Initial published results stem from high-end hardware including an M5 Max MacBook Pro, an iPhone 17 Pro, and a Galaxy S26 Ultra, with benchmarks for the AMD Ryzen AI Max+ 395 and Radeon 8060S marked as coming soon.

smartphone mobile device performance testing screen

Stock photo for illustration only, not from the actual event

Within Pipette, the core unit of measurement is a deployment configuration comprising model, quantization, runtime, and device. Benchmarks then evaluate specific metrics and token shapes to produce latency, throughput, or memory outcomes. Model quality is tracked separately via IFBench, GPQA Diamond, and MATH-500, with quality scores derived from llama.cpp evaluation runs on NVIDIA H100 80GB reference systems before being matched to on-device runs sharing identical models and quantization.

Never miss the latest news?

Subscribe to get news summaries by email - not often enough to be annoying.

โฆษณา

"On-device behavior is a property of the deployed system, not of the model in isolation."

Liquid AI

Separating model quality evaluation from actual on-device hardware runs by utilizing NVIDIA H100 reference systems represents a rigorous engineering approach. Because mobile chips and laptops face distinct thermal and power constraints compared to server GPUs, centralizing quality scoring ensures standardized accuracy before comparing those metrics against raw hardware performance speeds on consumer devices.

Performance runs follow a strict published methodology featuring fixed token shapes, greedy decoding, discarded warm-up periods, five measured repetitions, and readiness gating. Platform-specific checks verify thermal and load conditions prior to each timed repetition, and failing runs are omitted from publication. Furthermore, every submission records its benchmark version, token shape, model artifact, quantization, runtime settings, and underlying hardware.

Source: MarkTechPost

Comments

Leave a Comment
0/2000

Found something wrong in this article? Report an issue with this article