Ox Alpha Beats Claude and GPT on Benchmarks, Creator Unknown

Close-up of HTML code highlighted in vibrant colors on a computer monitor.

In brief

  • Ox Alpha appeared on OpenRouter August 20 with no company attribution
  • Free model outperforms Claude Fable 5 and GPT-5.6 Sol on coding benchmarks
  • Supports 1 million token context window, video input, and tool calling
  • Fingerprinting evidence suggests Zhipu AI's GLM family as likely source

The benchmark edge

Developer Ben Davis ran an initial 10-task DeepSWE sample and achieved 80% on Ox Alpha's first pass, compared to 65% for Claude Fable 5 and 52% for GPT-5.6 Sol. The model is described as a reasoning engine designed for coding, sustained agentic work, and production workloads.

Ox Alpha is free to use. It takes text, image, and video as input while returning text, and supports tool and function calling (though its JSON output isn't schema-enforced). The context window spans roughly 1 million tokens—substantial for long-form reasoning tasks.

The speed is striking. OpenCode claimed its route could handle 100 trillion tokens a day. Nous Research, offering the model free through its own portal, claimed capacity for 1 quadrillion tokens.

The identity puzzle

As of August 22, Ox Alpha had no entry on Artificial Analysis or LMSys Arena, the two main leaderboards developers rely on. Stripe CEO Patrick Collison called Ox Alpha "very impressive" on August 21, hours after his company agreed to acquire OpenRouter on August 19.

Developers floated Microsoft's unreleased MAI family, Xiaomi's MiMo line, DeepSeek, Alibaba's Qwen, and even Google as possible sources. But fingerprinting narrowed it down fast.

Independent analysis matched Ox Alpha's tokenizer to GLM-5.3 on every normalized test and its video-token spend to GLM-5V-Turbo. The model shares GLM's distinctive errors and audio-rejection behavior—Xiaomi's MiMo v2.5 accepts audio input, a headline feature Ox Alpha flatly rejects. DeepSeek has never shipped video capability and releases open weights instead of running stealth previews. Google and Qwen both use different tokenizers and video encoders entirely.

What's left points to Zhipu AI's family of models. The Beijing-based lab hasn't confirmed anything, but the technical fingerprints align too closely to ignore.

Nobody's talking. That's the story.