Tech Meridian ← LIVE FEED
RU

NEWS · MODELS · #435

Mistral launches Voxtral TTS, a 4B multilingual text-to-speech model

Mistral AI released Voxtral TTS, a 4-billion-parameter multilingual text-to-speech model that, according to Mistral, produces emotionally expressive, low-latency speech in nine languages and supports easy voice adaptation and zero-shot cross-lingual voice transfer. The model is available via API and Mistral Studio, is priced from $0.016 per 1K characters, and Mistral reports human-evaluation advantages versus ElevenLabs Flash v2.5 and parity with ElevenLabs v3 on quality while maintaining similar time-to-first-audio.

KEY POINTS

  1. Mistral AI released Voxtral TTS, a 4-billion-parameter multilingual text-to-speech model that, according to Mistral, produces emotionally expressive, low-latency speech in nine languages and supports easy voice adaptation and zero-shot cross-lingual voice transfer.
  2. The model is available via API and Mistral Studio, is priced from $0.016 per 1K characters, and Mistral reports human-evaluation advantages versus ElevenLabs Flash v2.5 and parity with ElevenLabs v3 on quality while maintaining similar time-to-first-audio.
  3. A compact, low-cost TTS model with fast voice adaptation and zero-shot cross-lingual transfer could materially shift options for enterprise voice agents and competing TTS providers.

WHY IT MATTERS

A compact, low-cost TTS model with fast voice adaptation and zero-shot cross-lingual transfer could materially shift options for enterprise voice agents and competing TTS providers.

SOURCES & TIMELINE

1