Qwen3-TTS
Qwen3-TTS is a family of advanced multilingual, controllable, robust, and streaming text-to-speech models. Trained on over 5 million hours of speech data spanning 10 languages, Qwen3-TTS supports state-of-the-art 3-second voice cloning and description-based control.
Apache-2.0
Text-to-Speech
PyTorch
Safetensors
English
Sign in to see model files
orCreate an account