2026•Tongyi Lab (Alibaba Group)
Qwen-Audio-3.0-TTS v3.0
Paid
Text-to-Speech
Multilingual
Qwen-Audio-3.0-TTS is a high-performance text-to-speech system from Alibaba's Tongyi Lab that supports 16 languages and various Chinese dialects. The model offers two primary variants, Flash and Plus, designed for real-time interaction and high-fidelity generation respectively, while featuring natural-language style control and robust voice cloning.
Support for 16 languages and 20 Chinese dialect regions
Dual-model architecture: Flash (low latency) and Plus (high-fidelity)
Natural-language style control for emotion, pace, and scenario