Liquid AI Releases LFM2.5-VL-3B: A 3B Vision-Language Model Built for On-Device Use
Liquid AI introduces the new 3-billion parameter vision-language model capable of screen reading, object grounding, and tool calling directly on devices.

Stock photo for illustration only, not from the actual event
- Liquid AI has launched the LFM2.5-VL-3B model featuring 3 billion parameters
- Built on the LFM2.5-2.6B language backbone and SigLIP2 NaFlex 400M vision encoder
- Supports a 32,768 token context length across 16 different languages
- Achieves an average score of 69.4 across 28 vision benchmarks
Liquid AI has officially rolled out its latest artificial intelligence model, LFM2.5-VL-3B, a 3-billion parameter vision-language model engineered to read screens, ground objects, and execute tool calls directly on devices. The checkpoint ships in four distinct formats including native, GGUF, ONNX, and MLX. Day-one runtimes feature llama.cpp, MLX, vLLM, SGLang, and ONNX, while requiring approximately 3 GB of memory to operate smoothly.
The architecture of LFM2.5-VL-3B expands upon the foundational design of LFM2-VL-3B along four specific axes. Its language backbone utilizes the LFM2.5-2.6B model, paired with a shape-optimized SigLIP2 NaFlex 400M vision encoder. The NaFlex framework expertly handles native resolutions by segmenting large images into non-overlapping 512x512 patches alongside a resized whole-image thumbnail. Furthermore, the model accommodates a context length of 32,768 tokens and offers robust support for 16 languages.

Stock photo for illustration only, not from the actual event
During the pre-training phase, the system ingested roughly 34 trillion tokens. Its vocabulary was doubled to 128,000 by extending the existing tokenizer in place, an enhancement that significantly improves non-Latin script coverage. Vision pre-training was scaled up by a factor of 4 in tokens using a curated and synthetic mix of captions, OCR, grounding, and instruction-following datasets. Subsequent post-training involved Supervised Fine-Tuning (SFT) leveraging knowledge distillation from a larger teacher model, Antidoom training, and multi-reward reinforcement learning.
The model follows a non-reasoning design philosophy, generating direct answers to maintain an optimal latency profile. When evaluated across 28 vision benchmarks utilizing vLLM 0.26.0 in non-reasoning mode, LFM2.5-VL-3B achieved an average score of 69.4. This performance matches the InternVL-3.5-4B model and lands a mere 0.7 points behind Qwen3.5-4B, both of which feature significantly larger 4.7-billion parameter architectures.
The introduction of LFM2.5-VL-3B highlights the ongoing industry shift toward highly efficient, small-scale AI models capable of running locally on edge devices. By packing multimodal vision and language capabilities into a roughly 3 GB footprint, Liquid AI enables developers to embed advanced intelligence directly into client applications without continuous cloud dependence, thereby lowering server costs and enhancing user data privacy.
Notable individual benchmark achievements include 73.1 on RealWorldQA compared to InternVL-3.5-4B at 67.7, and 84.3 on TextVQA against Qwen3.5-4B at 81.2. Additional scores feature 63.3 on MMStar, 68.5 on MathVista-mini, 81.3 on ChartQA, 91.1 on DocVQA, and 84.2 on OCRBench v1. However, CountBenchQA saw a slight regression to 87.3 from 92.2 in the previous iteration. On text-only evaluations, IFEval reached 82.3, improving from 72.9, while Gemma-4-E4B maintains the lead in that category with 87.9.
Source: MarkTechPost
Found something wrong in this article? Report an issue with this article
Comments
Leave a Comment