NEWS · MODELS · #435
Mistral launches Voxtral TTS, a 4B multilingual text-to-speech model
Mistral AI released Voxtral TTS, a 4-billion-parameter multilingual text-to-speech model that, according to Mistral, produces emotionally expressive, low-latency speech in nine languages and supports easy voice adaptation and zero-shot cross-lingual voice transfer. The model is available via API and Mistral Studio, is priced from $0.016 per 1K characters, and Mistral reports human-evaluation advantages versus ElevenLabs Flash v2.5 and parity with ElevenLabs v3 on quality while maintaining similar time-to-first-audio.
KEY POINTS
- Mistral AI released Voxtral TTS, a 4-billion-parameter multilingual text-to-speech model that, according to Mistral, produces emotionally expressive, low-latency speech in nine languages and supports easy voice adaptation and zero-shot cross-lingual voice transfer.
- The model is available via API and Mistral Studio, is priced from $0.016 per 1K characters, and Mistral reports human-evaluation advantages versus ElevenLabs Flash v2.5 and parity with ElevenLabs v3 on quality while maintaining similar time-to-first-audio.
- A compact, low-cost TTS model with fast voice adaptation and zero-shot cross-lingual transfer could materially shift options for enterprise voice agents and competing TTS providers.
WHY IT MATTERS
A compact, low-cost TTS model with fast voice adaptation and zero-shot cross-lingual transfer could materially shift options for enterprise voice agents and competing TTS providers.