OpenBMB Releases MiniCPM5-2B Model for Edge Devices
OpenBMB has launched MiniCPM5-2B, a 2.52B dense AI model scoring an average of 53.9 across 34 benchmarks and built for on-device deployment.

Stock photo for illustration only, not from the actual event
- OpenBMB introduces MiniCPM5-2B, a dense model featuring 2.52 billion parameters
- Achieves an average score of 53.9 across 34 benchmark evaluations
- Excels in tool use and code reasoning while operating directly on-device
- Released under the Apache 2.0 license alongside intermediate checkpoints and datasets
OpenBMB has officially released MiniCPM5-2B, a dense artificial intelligence model featuring 2.52 billion parameters. The model weights are distributed under an Apache 2.0 license and are designed to run smoothly through inference frameworks such as vLLM, SGLang, Transformers, llama.cpp, Ollama, LM Studio, MLX, and FlagOS, enabling efficient execution directly on local hardware and edge devices.
In benchmark evaluations, OpenBMB compared MiniCPM5-2B against similar-sized models including LFM2.5-2.6B, Qwen3.5-2B, and Gemma-4-E2B-it, alongside reference models like Qwen3.5-4B, granite-4.2-3B, Nemotron-3-Nano-4B, Gemma-4-E4B-it, and LFM2.5-8B-A1B. Across 34 evaluation benchmarks, MiniCPM5-2B achieved an overall average score of 53.9, outperforming other baseline models in its specific size class.
Looking into specific capabilities, the model scores 69.1 on LiveCodeBench v6 for code reasoning and 46.4 on SWE-bench Verified. Tool use represents its widest performance margin, recording 97.1 on τ²-Bench Telecom, 66.6 on BFCL v4, and 20.8 on τ³-Bench Banking. Long-context performance reaches 68.1 on NoLiMa, while general knowledge remains constrained by its size, posting 70.8 on MMLU-Pro and 8.9 on Humanity’s Last Exam.

Stock photo for illustration only, not from the actual event
The training pipeline follows the UltraData tiered data management method outlined in original research, passing through stable and decay base training phases followed by mid-training adaptation. Post-training incorporates 400 billion tokens of deep-thinking SFT and specialized RL teachers trained via the JustRL II algorithm for math, code, agents, and writing. The final phase applies on-policy distillation (OPD), merging 16 RL experts into a single model to boost reasoning and agentic benchmark scores.
"MiniCPM5-2B is a credible on-device option for agentic and tool-calling workloads, not a general knowledge model."
OpenBMB
OpenBMB's decision to publish intermediate checkpoints—including Base, Midtrain, and SFT-only variants—alongside open datasets like Ultra-FineWeb and UltraData-RL-2609 provides exceptional transparency. This allows developers to independently verify the exact contribution of each training stage rather than relying solely on aggregate headline metrics.
Alongside the model weights, OpenBMB published comprehensive resources including Ultra-FineWeb, Ultra-FineWeb-L3, UltraX, UltraData-Code, UltraData-Math, UltraData-SFT-2605, UltraData-SFT-Agent-2609 containing 500K agent samples, and UltraData-RL-2609 featuring over 80K RL samples to support further research and community development.
Source: MarkTechPost
Found something wrong in this article? Report an issue with this article
Comments
Leave a Comment