RELEASE · MODELS · #1429
Microsoft releases MAI-Transcribe-2-Streaming and MAI-Voice-2.1 speech models
Microsoft released MAI-Transcribe-2-Streaming, a real-time transcription model Microsoft says ranks first for accuracy on Artificial Analysis, supports 60 languages, and delivers first partial results in just over 100 ms; introductory pricing is $0.54 per hour of audio. The company also released MAI-Voice-2.1 and a low-latency variant MAI-Voice-2.1-Flash (about 150 ms) that speak 23 languages, can clone a voice from a few seconds of audio, have built-in safeguards, and are available via Microsoft Foundry, the MAI Playground and OpenRouter (Flash priced at $15 per million characters vs $22 for the standard plan, per Microsoft).
KEY POINTS
- Microsoft released MAI-Transcribe-2-Streaming, a real-time transcription model Microsoft says ranks first for accuracy on Artificial Analysis, supports 60 languages, and delivers first partial results in just over 100 ms; introductory pricing is $0.54 per hour of audio.
- The company also released MAI-Voice-2.1 and a low-latency variant MAI-Voice-2.1-Flash (about 150 ms) that speak 23 languages, can clone a voice from a few seconds of audio, have built-in safeguards, and are available via Microsoft Foundry, the MAI Playground and OpenRouter (Flash priced at $15 per million characters vs $22 for the standard plan, per Microsoft).
- Lower latency, multilingual coverage, voice-cloning from seconds of audio, and lower introductory pricing make these models likely to accelerate responsive voice agents and conversational experiences.
WHY IT MATTERS
Lower latency, multilingual coverage, voice-cloning from seconds of audio, and lower introductory pricing make these models likely to accelerate responsive voice agents and conversational experiences.