Topic: #benchmarking
-
MiniCPM5-2B tops sub-4B open models on Artificial Analysis Index v4.2
OpenBMB's MiniCPM5-2B, a 2-billion-parameter model built with Tsinghua University, claimed the top spot among open models under 4 billion parameters on Artificial Analysis Intelligence Index v4.2, signaling a shift toward efficient AI for edge devices.
-
CME Group and CF Benchmarks Launch Two New Multi-Asset Crypto Indices
CME Group and CF Benchmarks launched two new multi-asset digital asset indices on August 31, 2026. The CME CF Crypto Market Index tracks Bitcoin and Ether alongside broader holdings, while the CME CF Emerging Crypto Index excludes both, focusing on smaller digital assets. Both update every second and are designed for performance measurement and risk management rather than derivatives settlement.
-
Huawei's Claw-Anything benchmark exposes AI agent reliability gap
Researchers from Huawei and Chinese universities built a benchmark that simulates three months of real digital life, asking AI agents to complete tasks across multiple services and devices. GPT-5.5 scored just 34.5%—and fine-tuned open models outperformed some closed-source rivals.