Grok 4.7 ties MiMo-V2.6-Pro for first on Artificial Analysis Cyber Index
In brief
- Grok 4.7 and MiMo-V2.6-Pro each scored 56 to tie for first on the Cyber Index.
- GPT-6 Luna placed third on the Cyber Index with a score of 53.
- Grok 4.7 posted 68% pass@1 on CWE-Bench-AA and 74% on CyberGym-E2E-AA.
- Crypto Briefing called Grok 4.7's $11.67 per-task cost high next to cheaper models.
What the index tests
The Artificial Analysis Cyber Index tests AI agents on a set of enterprise cyber defense tasks. Models have to find and reproduce vulnerabilities inside real codebases, then ship working patches.
Artificial Analysis paired the launch with the Artificial Analysis Cyber Index Alliance, a group that includes IBM and NVIDIA, Crypto Briefing reported.
How Grok 4.7 scored
All of the figures below come from Artificial Analysis's Cyber Index, as reported by Crypto Briefing. On the CWE-Bench-AA sub-benchmark, Grok 4.7 recorded a 68% pass@1 rate, which tied for the lead on that test. Pass@1 measures whether a model gets the task right on its first attempt. On the CyberGym-E2E-AA patching task, the model hit a 74% success rate.
GPT-6 Luna landed third at 53. That's three points back.
The scores don't come cheap, though. Grok 4.7 costs $11.67 per task on the index, a figure Crypto Briefing described as high next to cheaper competing models.
A tie leaves cost in play
Grok 4.7 comes from SpaceXAI (the entity formed after the merger with xAI), according to Crypto Briefing. It follows Grok 4.6, which was benchmarked earlier in September 2026. Elon Musk highlighted the ranking on social media after the results came out, the outlet reported.
So what does a shared score of 56 tell enterprise buyers? Crypto Briefing's read is that Grok 4.7 and MiMo-V2.6-Pro are performing at a similar level on this test. In its view, cost and integration (along with reliability) would likely decide which model enterprises actually deploy.
It's a benchmark win on paper. The outlet's analysis suggests it isn't the whole buying decision.
Frequently asked questions
What does the Artificial Analysis Cyber Index measure?
The Cyber Index tests AI agents on a set of enterprise cyber defense tasks. Models have to find vulnerabilities, reproduce them, and ship working patches inside real codebases. It launched on or around September 28, 2026, according to Crypto Briefing.
What does pass@1 mean in the Grok 4.7 results?
Pass@1 measures whether a model gets a task right on its first attempt. Grok 4.7 recorded a 68% pass@1 rate on the CWE-Bench-AA sub-benchmark, which tied for the lead on that test, as reported by Crypto Briefing.
How much does Grok 4.7 cost per task on the Cyber Index?
Grok 4.7 costs $11.67 per task on the index, according to Crypto Briefing. The outlet described that figure as high next to cheaper competing models, and said cost, integration and reliability would likely decide which tied model enterprises deploy.


