Grok Voice Think Fast 2.0 tops Artificial Analysis Speech-to-Speech Index
In brief
- Grok Voice Think Fast 2.0 ranked first on Artificial Analysis Speech-to-Speech Index with 79.0% score
- Model achieved 97% speech reasoning, 94.7% task success, and 0.70-second time-to-first-audio
- xAI pricing: $0.08 per minute of audio, or $0.80 for a 10-minute call
- Supports 25+ languages and uses 60% fewer reasoning tokens than version 1.0
Performance metrics
Grok Voice Think Fast 2.0 achieved 97% for speech reasoning and 94.7% in task success rates. The model posts a time-to-first-audio of 0.70 seconds, enabling near-instantaneous responses in live conversations. An agentic performance score of 56.5% reflects the model's ability to handle multi-step tasks during dialogue—things like booking a reservation while answering follow-up questions about availability.
The efficiency gains are striking. Grok Voice Think Fast 2.0 uses roughly 60% fewer reasoning tokens than its predecessor, which launched approximately in April 2026. That jump came less than three months after the 1.0 release, signaling rapid iteration in xAI's voice AI roadmap.
Pricing and deployment
xAI is charging $0.08 per minute of audio, which puts a 10-minute customer service call at $0.80 in AI costs. The company has set up an automatic upgrade path for developers using the grok-voice-latest API endpoint, with the migration initiated on August 5, 2026.
Competitive position
The model competes against offerings from Qwen and other major players in the voice AI space. While Grok Voice Think Fast 2.0 trails slightly in some individual metrics, its overall index placement at the top suggests it strikes the best balance across the full range of evaluated capabilities.
The 2.0 version supports over 25 languages and has improved handling of noisy environments and various accents. These enhancements broaden the model's appeal for global deployment in customer service, accessibility, and conversational AI applications.
"Agentic performance measures how well a voice model can handle multi-step tasks during a conversation, things like booking a reservation while answering follow-up questions about availability." — Crypto Briefing reporting


