MiniCPM5-2B tops sub-4B open models on Artificial Analysis Index v4.2

Editorial illustration: A silver phone-like device with a black screen and exposed turquoise chip rests on a stone pedestal between two stacks of larger dark processors.

In brief

  • MiniCPM5-2B scored 15 points on Artificial Analysis v4.2, topping sub-4B open models
  • Model runs on phones, laptops, and devices with limited computational resources
  • OpenBMB optimized for AMD, Intel, MediaTek, and Qualcomm chips
  • Dense transformer architecture keeps all parameters active during inference
  • Independent performance validation not yet publicly documented

Efficient models gain traction

The ranking reflects a broader shift in AI development. Artificial Analysis released version 4.2 of its Intelligence Index on September 4, 2026, introducing new methodology that now weights private test sets at 40% of total scoring, up from previous versions. The update added the AA-Briefcase agentic evaluation and Surge's GDP.pdf long-context test, reflecting how benchmarking itself is evolving.

MiniCPM5-2B is a dense transformer, meaning every parameter is active during inference rather than routing through a mixture-of-experts architecture. This design choice affects both performance and resource consumption. The model supports context windows ranging from 131K to 512K tokens depending on configuration, and OpenBMB has optimized it for chips from AMD, Intel, MediaTek, and Qualcomm.

Performance claims and validation

OpenBMB claims state-of-the-art results. The team says the model delivers state-of-the-art performance within its parameter range across coding, mathematics, long-context understanding, tool utilization, and agentic workflows. Training incorporated reinforcement learning alignment and what OpenBMB describes as high-quality trajectories.

The numbers merit scrutiny. MiniCPM5-2B's predecessor, MiniCPM5-1B, previously scored 17.9 on an earlier version of the Artificial Analysis Index, claiming the top position among models under 2 billion parameters. The score drop from 17.9 to 15 reflects methodology changes rather than performance decline — the scoring difference between the two models reflects changes in index methodology.

Competitive landscape

MiniCPM5-2B operates in crowded territory. Meta's Llama series, Microsoft's Phi models, and Google's Gemma variants all compete in similar parameter ranges. Meanwhile, Anthropic's Claude Fable 5.1 leads the overall leaderboard.

The critical gap: independent validation of the model's claimed performance has not yet been publicly documented. In benchmarking, third-party reproduction separates capability claims from press releases. OpenBMB's rankings matter, but they're not yet independently verified by the broader research community.