Skip to main content

OpenBMB Releases MiniCPM5-2B Model for Edge Devices

OpenBMB has launched MiniCPM5-2B, a 2.52B dense AI model scoring an average of 53.9 across 34 benchmarks and built for on-device deployment.

AI-written
Inewgen
08 Sep 2026Source: MarkTechPost3 min read (0 views)
Share
OpenBMB Releases MiniCPM5-2B Model for Edge Devices

Stock photo for illustration only, not from the actual event

Font size
  • OpenBMB introduces MiniCPM5-2B, a dense model featuring 2.52 billion parameters
  • Achieves an average score of 53.9 across 34 benchmark evaluations
  • Excels in tool use and code reasoning while operating directly on-device
  • Released under the Apache 2.0 license alongside intermediate checkpoints and datasets

OpenBMB has officially released MiniCPM5-2B, a dense artificial intelligence model featuring 2.52 billion parameters. The model weights are distributed under an Apache 2.0 license and are designed to run smoothly through inference frameworks such as vLLM, SGLang, Transformers, llama.cpp, Ollama, LM Studio, MLX, and FlagOS, enabling efficient execution directly on local hardware and edge devices.

In benchmark evaluations, OpenBMB compared MiniCPM5-2B against similar-sized models including LFM2.5-2.6B, Qwen3.5-2B, and Gemma-4-E2B-it, alongside reference models like Qwen3.5-4B, granite-4.2-3B, Nemotron-3-Nano-4B, Gemma-4-E4B-it, and LFM2.5-8B-A1B. Across 34 evaluation benchmarks, MiniCPM5-2B achieved an overall average score of 53.9, outperforming other baseline models in its specific size class.

53.9Average across 34 benchmarks
2.52BModel parameter count
97.1Score on τ²-Bench Telecom

Looking into specific capabilities, the model scores 69.1 on LiveCodeBench v6 for code reasoning and 46.4 on SWE-bench Verified. Tool use represents its widest performance margin, recording 97.1 on τ²-Bench Telecom, 66.6 on BFCL v4, and 20.8 on τ³-Bench Banking. Long-context performance reaches 68.1 on NoLiMa, while general knowledge remains constrained by its size, posting 70.8 on MMLU-Pro and 8.9 on Humanity’s Last Exam.

chromebook notebook computer office desk workspace

Stock photo for illustration only, not from the actual event

The training pipeline follows the UltraData tiered data management method outlined in original research, passing through stable and decay base training phases followed by mid-training adaptation. Post-training incorporates 400 billion tokens of deep-thinking SFT and specialized RL teachers trained via the JustRL II algorithm for math, code, agents, and writing. The final phase applies on-policy distillation (OPD), merging 16 RL experts into a single model to boost reasoning and agentic benchmark scores.

Never miss the latest news?

Subscribe to get news summaries by email - not often enough to be annoying.

โฆษณา

"MiniCPM5-2B is a credible on-device option for agentic and tool-calling workloads, not a general knowledge model."

OpenBMB

OpenBMB's decision to publish intermediate checkpoints—including Base, Midtrain, and SFT-only variants—alongside open datasets like Ultra-FineWeb and UltraData-RL-2609 provides exceptional transparency. This allows developers to independently verify the exact contribution of each training stage rather than relying solely on aggregate headline metrics.

Alongside the model weights, OpenBMB published comprehensive resources including Ultra-FineWeb, Ultra-FineWeb-L3, UltraX, UltraData-Code, UltraData-Math, UltraData-SFT-2605, UltraData-SFT-Agent-2609 containing 500K agent samples, and UltraData-RL-2609 featuring over 80K RL samples to support further research and community development.

Source: MarkTechPost

Comments

Leave a Comment
0/2000

Found something wrong in this article? Report an issue with this article