Not established from the available sources.
ALIBABA (QWEN) · MODEL RELEASE TRACKER
Qwen-Audio-3.1
Qwen-Audio-3.1 is a released lineup of five audio models from Alibaba's Qwen team for automatic speech recognition (ASR), text-to-speech (TTS), and real-time interaction, introducing new features and substantial price cuts.CURRENT SNAPSHOT0/5 DIMENSIONS WITH DATA
The dimensions that change the decision.
Not established from the available sources.
Not established from the available sources.
Not established from the available sources.
Not established from the available sources.
RELEASE TIMELINE
Published, source-backed release events only.
VERIFIABLE FACTS
Every value stays attached to a source and date.
RELEASE · Release / announcementDEVELOPER CLAIM
Qwen-Audio-3.1 was released/announced on 2026-09-23 (per the published article).
CAPABILITIES · Model lineup and modalitiesDEVELOPER CLAIM
The release is a lineup of five models covering speech recognition (ASR), text-to-speech (TTS), and real-time interaction.
CAPABILITIES · ASR improvementsDEVELOPER CLAIM
The ASR model improves multilingual and dialect recognition and automatically cleans up filler words and repetitions.
CAPABILITIES · ASR-Next featuresDEVELOPER CLAIM
ASR-Next adds multi-speaker identification with timestamps and detects emotions, ambient sounds, and machine noise.
CAPABILITIES · TTS capabilitiesDEVELOPER CLAIM
TTS supports multilingual synthesis with natural cross-language voice transfer; users can control emotion, speed, and style via text prompts.
CAPABILITIES · TTS-Next approachDEVELOPER CLAIM
TTS-Next pairs a language model with a diffusion approach to generate voice, sound effects, and background audio in a single pass.
CAPABILITIES · Real-time interaction featuresDEVELOPER CLAIM
The real-time model supports simultaneous speaking and listening with instant interruption.
SAFETY · Adaptive response behaviorDEVELOPER CLAIM
The real-time model reportedly detects low mood and responds more slowly and with more empathy.
WHAT CHANGED