Qwen3-TTSAlibaba
|
Simba 3.2Speechify
|
|||||
Related Products
|
||||||
About
Qwen3-TTS is an open source series of advanced text-to-speech models developed by the Qwen team at Alibaba Cloud under the Apache-2.0 license, offering stable, expressive, and real-time speech generation with features such as voice cloning, voice design, and fine-grained control of prosody and acoustic attributes. The models support 10 major languages, including Chinese, English, Japanese, Korean, German, French, Russian, Portuguese, Spanish, and Italian, and multiple dialectal voice profiles with adaptive control over tone, speaking rate, and emotional expression based on text semantics and instructions. Qwen3-TTS uses efficient tokenization and a dual-track architecture that enables ultra-low-latency streaming synthesis (first audio packet in ~97 ms), making it suitable for interactive and real-time use cases, and includes a range of models with different capabilities (e.g., rapid 3-second voice cloning, custom voice timbres, and instruction-based voice design).
|
About
Speechify’s text-to-speech API offers a family of Simba models for real-time voice generation across English, European languages, and broader multilingual use cases. Simba 3.2 is recommended for new English integrations, providing streaming-native synthesis, the lowest time to first byte, richer expressivity than earlier generations, and full support for SSML and emotion control. Simba 3.0 extends streaming-native speech to English, German, Spanish, French, Italian, and Brazilian Portuguese, with language selection handled through the request or voice locale. Simba Multilingual supports 35 locales across 30 languages, including mixed-language content and automatic language detection, while Simba English remains available as a legacy model for compatibility. Developers select a model through one parameter and can switch without changing the rest of the request structure, including voice, format, and SSML settings.
|
|||||
Platforms Supported
Windows
Not Supported
Mac
Not Supported
Linux
Not Supported
Cloud
Supported
On-Premises
Not Supported
iPhone
Not Supported
iPad
Not Supported
Android
Not Supported
Chromebook
Not Supported
|
Platforms Supported
Windows
Not Supported
Mac
Not Supported
Linux
Not Supported
Cloud
Supported
On-Premises
Not Supported
iPhone
Not Supported
iPad
Not Supported
Android
Not Supported
Chromebook
Not Supported
|
|||||
Audience
Researchers who need a model for expressive, multilingual, controllable, and streaming voice generation in applications like voice assistants, dubbing, accessibility, and creative audio synthesis
|
Audience
Developers that need to generate expressive, low-latency multilingual speech and cloned voices for voice-enabled applications
|
|||||
Support
Phone Support
Not Supported
24/7 Live Support
Not Supported
Online
Supported
|
Support
Phone Support
Not Supported
24/7 Live Support
Not Supported
Online
Supported
|
|||||
API
Offers API
Supported
|
API
Offers API
Supported
|
|||||
Screenshots and Videos |
Screenshots and Videos |
|||||
Pricing
Free
Free Version
Supported
Free Trial
Not Supported
|
Pricing
No information available.
Free Version
Not Supported
Free Trial
Not Supported
|
|||||
Reviews/
|
Reviews/
|
|||||
Training
Documentation
Supported
Webinars
Not Supported
Live Online
Supported
In Person
Not Supported
|
Training
Documentation
Supported
Webinars
Not Supported
Live Online
Not Supported
In Person
Not Supported
|
|||||
Company InformationAlibaba
Founded: 1999
China
github.com/QwenLM/Qwen3-TTS
|
Company InformationSpeechify
Founded: 2017
United States
docs.speechify.ai/build/guides/concepts/models
|
|||||
Alternatives |
Alternatives |
|||||
|
|
|
|||||
|
|
||||||
|
|
|
|||||
|
|
|
|||||
Categories |
Categories |
|||||
Integrations
Alibaba Cloud
Not Supported
OpenClaw
Not Supported
Qwen
Not Supported
|
||||||
|
|
|