PricingPricing starts at $0.016 per 1,000 characters.
MISTRAL AI · MODEL RELEASE TRACKER
Voxtral TTS API
Voxtral TTS is a 4B-parameter text-to-speech model announced by Mistral AI on 2026-03-23. It produces realistic, emotionally expressive multilingual speech (9 languages), supports speaker modeling and zero-shot cross-lingual voice adaptation, targets enterprise voice workflows and scalable agents, and is offered via API and Mistral Studio. Pricing starts at $0.016 per 1,000 characters.CURRENT SNAPSHOT2/5 DIMENSIONS WITH DATA
The dimensions that change the decision.
Not established from the available sources.
ModalityText-to-speech (TTS) model.
Not established from the available sources.
Not established from the available sources.
VERIFIABLE FACTS
Every value stays attached to a source and date.
RELEASE · Launch dateDEVELOPER CLAIM
Announced/launched on 2026-03-23.
MODALITIES · ModalityDEVELOPER CLAIM
Text-to-speech (TTS) model.
CAPABILITIES · Model sizeDEVELOPER CLAIM
4 billion parameters (described as lightweight).
CAPABILITIES · LanguagesDEVELOPER CLAIM
Realistic, emotionally expressive speech in 9 languages with support for diverse dialects.
CAPABILITIES · Expressiveness and speaker modelingDEVELOPER CLAIM
Excels at contextual understanding (e.g., neutral, happy, sarcastic) and speaker modeling, capturing personality, pauses, rhythm, intonation, and emotional dexterity.
CAPABILITIES · Cross-lingual voice adaptationDEVELOPER CLAIM
Supports zero-shot cross-lingual voice adaptation (adapting a voice across languages without additional training).
CAPABILITIES · Latency and scalabilityDEVELOPER CLAIM
Described as low latency and cost-effective, intended for enterprise-grade voice workflows and scalable AI agents.
PRICING · PricingDEVELOPER CLAIM
Pricing starts at $0.016 per 1,000 characters.
WHAT CHANGED