Scoop
AI tool news · rumor vs. reality

The AI Wire ●

Rumors tracked. Announcements verified. Updated daily.

RELEASED

xAI ships Grok Voice Transcribe 2.0 with doubled accuracy claims, same pricing

SpaceXAI (xAI) released Grok Voice Transcribe 2.0 on September 18, a speech-to-text model it calls the most accurate available — claiming double the accuracy of version 1.0 at unchanged pricing: $0.10 per audio hour for batch transcription and $0.20 for real-time streaming, with speaker diarization, word-level timestamps, and key-term biasing included. It is built on the audio foundation behind Grok Voice, trained on live, noisy, multilingual audio and refined with post-training.

The company says word error rate fell from 20.6% to 6.8% on short multilingual voice commands, from 10.6% to 7.1% on telephony, and from 8.7% to 3.3% on conversational audio — and that it ranks first among 32 streaming models on the public Artificial Analysis leaderboard. Those are vendor-reported numbers; independent confirmation matters more than usual here, since xAI's Terminal-Bench claims on Grok 4.7 recently showed a 12-point gap to independent runs.

Distribution is the interesting part: Atlassian's Loom now uses it to transcribe every video, per RuntimeWire. The model is live through the Speech-to-Text API as grok-voice-transcribe-2.0, with 1.0 headed for deprecation in the coming weeks — developers who need the old model should pin grok-voice-transcribe-1.0 explicitly.

Sources

← Back to headlines