RELEASE · MODELS · #860
Alibaba releases Qwen-Audio-3.1 suite and cuts AI audio prices up to 95%
Alibaba's Qwen team released Qwen-Audio-3.1, a set of five models for automatic speech recognition (ASR), text‑to‑speech (TTS) and real‑time interaction. New capabilities include improved multilingual and dialect ASR, multi‑speaker timestamps and emotion/ambient/machine‑noise detection in ASR-Next, cross‑language and controllable TTS and a diffusion‑based TTS-Next that generates voice, effects and background audio in one pass; the real‑time model supports simultaneous speaking/listening with instant interruption and mood‑aware responses. Alibaba also sharply cut prices on its audio services (TTS ≈70% lower, Realtime ≈85% lower, ASR up to 95%) and made details available via its blog and Qwen Cloud.
KEY POINTS
- Alibaba's Qwen team released Qwen-Audio-3.1, a set of five models for automatic speech recognition (ASR), text‑to‑speech (TTS) and real‑time interaction.
- New capabilities include improved multilingual and dialect ASR, multi‑speaker timestamps and emotion/ambient/machine‑noise detection in ASR-Next, cross‑language and controllable TTS and a diffusion‑based TTS-Next that generates voice, effects and background audio in one pass; the real‑time model supports simultaneous speaking/listening with instant interruption and mood‑aware responses.
- Alibaba also sharply cut prices on its audio services (TTS ≈70% lower, Realtime ≈85% lower, ASR up to 95%) and made details available via its blog and Qwen Cloud.
WHY IT MATTERS
The release adds advanced ASR/TTS and real‑time audio capabilities while steep price cuts could accelerate adoption of AI audio services and pressure competitors on cost.