Not established from the available sources.
MISTRAL AI · MODEL RELEASE TRACKER
Voxtral Realtime
Voxtral Realtime is a streaming speech-to-text model in the Voxtral Transcribe 2 family, purpose-built for live transcription with configurable latency (down to sub-200ms). It is released with open weights under the Apache 2.0 license and is optimized to run on edge devices (4B parameters).CURRENT SNAPSHOT3/5 DIMENSIONS WITH DATA
The dimensions that change the decision.
Not established from the available sources.
ModalitySpeech-to-text (streaming / live transcription).
Benchmarks and reported performanceEvaluated on the FLEURS transcription benchmark and multiple diarization benchmarks (Switchboard, CallHome, AMI-IHM, AMI-SDM, SBCSAE, TalkBank). Reported: at 2.4s delay Realtime matches Voxtral Mini Transcribe V2; at 480ms delay it stays within ~1–2% word error rate.
Weights and license / availabilityVoxtral Realtime ships with open weights under the Apache 2.0 license and the weights are released on the Hugging Face Hub; stated as deployable on edge.
VERIFIABLE FACTS
Every value stays attached to a source and date.
RELEASE · Release / announcement dateDEVELOPER CLAIM
Announced/released on 2026-02-04 as part of the Voxtral Transcribe 2 release.
MODALITIES · ModalityDEVELOPER CLAIM
Speech-to-text (streaming / live transcription).
CAPABILITIES · Configurable latency (streaming)DEVELOPER CLAIM
Purpose-built for live transcription with latency configurable down to sub-200ms; uses a streaming architecture that transcribes audio as it arrives.
CAPABILITIES · Transcription quality and diarizationDEVELOPER CLAIM
Described as delivering state-of-the-art transcription quality and diarization (family-level description applied to Voxtral Realtime in the release).
BENCHMARKS · Benchmarks and reported performanceDEVELOPER CLAIM
Evaluated on the FLEURS transcription benchmark and multiple diarization benchmarks (Switchboard, CallHome, AMI-IHM, AMI-SDM, SBCSAE, TalkBank). Reported: at 2.4s delay Realtime matches Voxtral Mini Transcribe V2; at 480ms delay it stays within ~1–2% word error rate.
CAPABILITIES · Multilingual supportDEVELOPER CLAIM
Natively multilingual with strong transcription performance in 13 languages (including English, Chinese, Hindi, Spanish, Arabic, French, Portuguese, Russian, German, Japanese, Korean, Italian, Dutch).
CAPABILITIES · Model size and edge efficiencyDEVELOPER CLAIM
Has a ~4B parameter footprint and is described as running efficiently on edge devices.
AVAILABILITY · Weights and license / availabilityDEVELOPER CLAIM
Voxtral Realtime ships with open weights under the Apache 2.0 license and the weights are released on the Hugging Face Hub; stated as deployable on edge.
WHAT CHANGED