Not established from the available sources.
MISTRAL AI · MODEL RELEASE TRACKER
Voxtral Transcribe 2
Voxtral Transcribe 2 is a released family of next-generation speech-to-text models from Mistral AI (Voxtral Mini Transcribe V2 for batch and Voxtral Realtime for live). It offers state-of-the-art transcription, speaker diarization, word-level timestamps, multilingual support (13 languages), and configurable ultra-low latency; Voxtral Realtime weights are released under the Apache 2.0 license on the Hugging Face Hub.CURRENT SNAPSHOT2/5 DIMENSIONS WITH DATA
The dimensions that change the decision.
Not established from the available sources.
ModalitySpeech-to-text (transcription) with speaker diarization and word-level timestamps.
FLEURS benchmark and pricing claimReported word error rate (WER) results across languages on the FLEURS transcription benchmark; Voxtral Mini Transcribe V2 is claimed to achieve the lowest WER at the lowest price point.
Realtime WER at different delaysAt 2.4 seconds delay, Realtime matches Voxtral Mini Transcribe V2; at 480 ms delay, Realtime 'stays within 1–2% word error rate' (as reported), enabling voice agents with near-offline accuracy.
Not established from the available sources.
VERIFIABLE FACTS
Every value stays attached to a source and date.
RELEASE · Release dateDEVELOPER CLAIM
Announced/released 2026-02-04 (publication date of the announcement).
RELEASE · Included models in the familyDEVELOPER CLAIM
Voxtral Mini Transcribe V2 (batch transcription) and Voxtral Realtime (live/streaming applications).
MODALITIES · ModalityDEVELOPER CLAIM
Speech-to-text (transcription) with speaker diarization and word-level timestamps.
CAPABILITIES · Languages supportedDEVELOPER CLAIM
Natively multilingual, supporting 13 languages: English, Chinese, Hindi, Spanish, Arabic, French, Portuguese, Russian, German, Japanese, Korean, Italian, and Dutch.
CAPABILITIES · Streaming architecture and latencyDEVELOPER CLAIM
Voxtral Realtime uses a novel streaming architecture that transcribes audio as it arrives; latency is configurable down to sub-200 ms (purpose-built for live transcription and voice agents).
CAPABILITIES · Model size / footprintDEVELOPER CLAIM
Reported 4B parameter footprint (described as running efficiently on edge devices).
BENCHMARKS · FLEURS benchmark and pricing claimDEVELOPER CLAIM
Reported word error rate (WER) results across languages on the FLEURS transcription benchmark; Voxtral Mini Transcribe V2 is claimed to achieve the lowest WER at the lowest price point.
BENCHMARKS · Realtime WER at different delaysDEVELOPER CLAIM
At 2.4 seconds delay, Realtime matches Voxtral Mini Transcribe V2; at 480 ms delay, Realtime 'stays within 1–2% word error rate' (as reported), enabling voice agents with near-offline accuracy.
WHAT CHANGED