xAI ships Grok Voice Think Fast 2.0
xAI's new speech-to-speech model cuts response latency to 0.70s and improves transcription accuracy, priced at $0.08 per audio minute, replacing v1 on August 5.
xAI released Grok Voice Think Fast 2.0, its next-generation speech-to-speech model, on July 29. Per xAI’s own benchmarking against Artificial Analysis, the new model scores 82.9% on the AA Speech-to-Speech Quality Index versus 75.7% for its predecessor, cuts time-to-first-audio from 1.25s to 0.70s, and improves transcription accuracy 1.5-2x over Deepgram Nova 3 and ElevenLabs Scribe v2 across 24 languages, a gap that widens to roughly 10x in noisy, telephony-compressed conditions. xAI says the model reasons through queries while speaking rather than adding latency to think first, and cites an A/B test on its Starlink support line (+1 888 GO STARLINK) showing a meaningful lift in sales conversion and support containment. Think Fast 2.0 is priced at $0.08 per minute of audio; the grok-voice-latest alias will automatically switch from Think Fast 1.0 to 2.0 on August 5, with no prompt changes required, or developers can pin grok-voice-think-fast-1.0 to stay on the older model.
What it means for operators
The concrete numbers here, sub-second first-audio latency and a real containment-rate lift on a live support line, are the kind of thing worth testing against your own voice-agent traffic before the August 5 auto-upgrade lands. If your prompts are tuned around v1’s slower, less accurate transcription, budget time to re-validate them rather than assuming the swap is purely additive.