PricingAnnouncement states v2.0 delivers twice the accuracy of Grok Voice Transcribe 1.0 at the same price.
SPACEXAI · MODEL RELEASE TRACKER
Grok Voice Transcribe 2.0
Grok Voice Transcribe 2.0 is SpaceXAI's released speech-to-text model (announced 2026-09-18). It claims improved real-world transcription accuracy versus Grok Voice Transcribe 1.0, multilingual automatic language detection and mid-recording language switching, leaderboard-leading accuracy on a public streaming benchmark, and availability via existing Speech-to-Text API (batch and streaming).CURRENT SNAPSHOT4/5 DIMENSIONS WITH DATA
The dimensions that change the decision.
Not established from the available sources.
ModalitySpeech-to-text (audio transcription) — an audio transcription model.
Public benchmark rankingRanks first for accuracy among 32 streaming models on the public Artificial Analysis leaderboard (as reported in the announcement).
Word error rate (WER) improvementsReportedly twice as accurate as Grok Voice Transcribe 1.0 overall; on a short-phrase set WER reportedly falls from 20.6% (v1.0) to 6.8% (v2.0); internal evaluations across four production-derived sets show improvements across all sets, with telephony leading every tested model.
Availability / integration modesAvailable via existing Speech-to-Text API integrations with no code changes; supports Batch and streaming modes and can transcribe recorded files, URLs, or live audio streams.
RELEASE TIMELINE
Published, source-backed release events only.
VERIFIABLE FACTS
Every value stays attached to a source and date.
RELEASE · Release/announcement dateDEVELOPER CLAIM
Announcement published 2026-09-18T17:34:03.777368+00:00 (model release announced by SpaceXAI).
MODALITIES · ModalityDEVELOPER CLAIM
Speech-to-text (audio transcription) — an audio transcription model.
CAPABILITIES · Multilingual support and language detectionDEVELOPER CLAIM
Transcribes dozens of languages, automatically detects language, and follows mid-recording language switches in a single pass.
CAPABILITIES · Foundation model and training dataDEVELOPER CLAIM
Built on the audio foundation model behind Grok Voice; trained on a dataset of live, noisy, multilingual audio across diverse environments and refined with post-training.
BENCHMARKS · Public benchmark rankingDEVELOPER CLAIM
Ranks first for accuracy among 32 streaming models on the public Artificial Analysis leaderboard (as reported in the announcement).
BENCHMARKS · Word error rate (WER) improvementsDEVELOPER CLAIM
Reportedly twice as accurate as Grok Voice Transcribe 1.0 overall; on a short-phrase set WER reportedly falls from 20.6% (v1.0) to 6.8% (v2.0); internal evaluations across four production-derived sets show improvements across all sets, with telephony leading every tested model.
PRICING · PricingDEVELOPER CLAIM
Announcement states v2.0 delivers twice the accuracy of Grok Voice Transcribe 1.0 at the same price.
AVAILABILITY · Availability / integration modesDEVELOPER CLAIM
Available via existing Speech-to-Text API integrations with no code changes; supports Batch and streaming modes and can transcribe recorded files, URLs, or live audio streams.
WHAT CHANGED