Not established from the available sources.
XAI · MODEL RELEASE TRACKER
grok-voice-latest (Grok Voice Think Fast 2.0)
Grok Voice Think Fast 2.0 — xAI's announced next-generation speech-to-speech voice model, claiming improved intelligence, transcription accuracy, conversational abilities, and efficient reasoning during speech.CURRENT SNAPSHOT2/5 DIMENSIONS WITH DATA
The dimensions that change the decision.
Not established from the available sources.
ModalityDescribed as a speech-to-speech voice model (speech in → speech out).
Transcription accuracy vs competitorsReported 1.5–2.0× improvement versus Deepgram Nova 3 and ElevenLabs Scribe v2 across thousands of short phrases in 24 languages; 1.4× improvement versus Grok Voice Think Fast 1.0. In noisy settings the gap is reported to widen to about 10×.
Not established from the available sources.
VERIFIABLE FACTS
Every value stays attached to a source and date.
RELEASE · Announcement dateDEVELOPER CLAIM
Announced in a press post published 2026-09-16.
MODALITIES · ModalityDEVELOPER CLAIM
Described as a speech-to-speech voice model (speech in → speech out).
BENCHMARKS · Transcription accuracy vs competitorsDEVELOPER CLAIM
Reported 1.5–2.0× improvement versus Deepgram Nova 3 and ElevenLabs Scribe v2 across thousands of short phrases in 24 languages; 1.4× improvement versus Grok Voice Think Fast 1.0. In noisy settings the gap is reported to widen to about 10×.
CAPABILITIES · Reasoning during speech and latencyDEVELOPER CLAIM
The model is said to 'reason through queries while speaking' (reasoning in parallel with speech) with no impact on latency; trained to be more efficient with reasoning tokens so tool calls are often executed before the end of the agent's first sentence.
CAPABILITIES · Conversational behaviorDEVELOPER CLAIM
Trained with extensive reinforcement learning to favor shorter sentences, ask one question at a time, and avoid fluff, aiming for simpler, more fluid conversations without edits to existing prompts.
WHAT CHANGED