Tech Meridian← ALL MODELS
PROMY MERIDIAN

SPACEXAI · MODEL RELEASE TRACKER

Grok Voice Transcribe 2.0

Grok Voice Transcribe 2.0 is SpaceXAI's released speech-to-text model (announced 2026-09-18). It claims improved real-world transcription accuracy versus Grok Voice Transcribe 1.0, multilingual automatic language detection and mid-recording language switching, leaderboard-leading accuracy on a public streaming benchmark, and availability via existing Speech-to-Text API (batch and streaming).

CURRENT SNAPSHOT4/5 DIMENSIONS WITH DATA

The dimensions that change the decision.

PRICING
PricingAnnouncement states v2.0 delivers twice the accuracy of Grok Voice Transcribe 1.0 at the same price.
CONTEXT WINDOW

Not established from the available sources.

MODALITIES
ModalitySpeech-to-text (audio transcription) — an audio transcription model.
BENCHMARKS
Public benchmark rankingRanks first for accuracy among 32 streaming models on the public Artificial Analysis leaderboard (as reported in the announcement).
Word error rate (WER) improvementsReportedly twice as accurate as Grok Voice Transcribe 1.0 overall; on a short-phrase set WER reportedly falls from 20.6% (v1.0) to 6.8% (v2.0); internal evaluations across four production-derived sets show improvements across all sets, with telephony leading every tested model.
AVAILABILITY
Availability / integration modesAvailable via existing Speech-to-Text API integrations with no code changes; supports Batch and streaming modes and can transcribe recorded files, URLs, or live audio streams.

RELEASE TIMELINE

Published, source-backed release events only.

1 SOURCES · IMPORTANCE 7.0

SpaceXAI releases Grok Voice Transcribe 2.0 speech-to-text model

SpaceXAI announced Grok Voice Transcribe 2.0, an updated speech-to-text model built on the Grok Voice audio foundation. The company says it is twice as accurate as Grok Voice Transcribe 1.0 in real-world tests, ranks first on the Artificial Analysis streaming leaderboard, supports dozens of languages with automatic detection and speaker diarization, and keeps the same pricing and API compatibility as 1.0 while the older version is phased out.

→

VERIFIABLE FACTS

Every value stays attached to a source and date.

CAPABILITIES · Foundation model and training dataDEVELOPER CLAIM

Built on the audio foundation model behind Grok Voice; trained on a dataset of live, noisy, multilingual audio across diverse environments and refined with post-training.

BENCHMARKS · Word error rate (WER) improvementsDEVELOPER CLAIM

Reportedly twice as accurate as Grok Voice Transcribe 1.0 overall; on a short-phrase set WER reportedly falls from 20.6% (v1.0) to 6.8% (v2.0); internal evaluations across four production-derived sets show improvements across all sets, with telephony leading every tested model.

AVAILABILITY · Availability / integration modesDEVELOPER CLAIM

Available via existing Speech-to-Text API integrations with no code changes; supports Batch and streaming modes and can transcribe recorded files, URLs, or live audio streams.

WHAT CHANGED

Stored passport versions, without reconstructed history.

Passport created8 facts