AI & Voice

ElevenLabs-Tier Voice Cloning, Now Running on Qwen-Audio-TTS

January 30, 2026 ·4 min read
ElevenLabs-Tier Voice Cloning, Now Running on Qwen-Audio-TTS
Updated August 2026. This post originally announced our move to Qwen3 TTS in January 2026. Voice Cloning now runs on Qwen-Audio-TTS, the current generation of the same family — Alibaba is retiring the older Qwen3-TTS-VC line in October 2026. The quality story below still holds; the model behind it is newer, faster and covers more languages.

At simpleTTS.ai, our goal has always been to make high-quality speech generation accessible and easy. Voice Cloning is where that goal is hardest to hit, and where your choice of model matters most.

Our Voice Cloning feature runs on Qwen-Audio-TTS, Alibaba's current-generation cloning model.

If you follow AI audio news, you already know why that matters. If you don't, here is the short version: in early 2026 the benchmark for realistic AI speech shifted, and Qwen's cloning models were what shifted it.

The Elephant in the Room: The ElevenLabs Comparison

For years there was one undisputed name in the AI voice space: ElevenLabs. They set the standard for emotional realism and prosody that most other models struggled to match.

In early 2026, that stopped being a foregone conclusion.

When Qwen3 TTS landed, developers, sound engineers and creators ran it side by side against the incumbent on X (Twitter), Reddit and YouTube. The consensus was striking: on speaker similarity (how much the clone sounds like the original person) and accent preservation, many testers put Qwen at or above the model everyone else was measured against.

That is why we moved Voice Cloning onto Qwen's cloning line, and why we have stayed on it through this year's model refresh.

What This Means for Your Projects

Running Voice Cloning on Qwen-Audio-TTS means tangible things for your projects:

1. Unprecedented Realism with Less Data

Forget training for hours. Qwen's cloning models are "zero-shot": a clean 10-to-30-second clip of the target voice is enough to produce a clone that captures unique vocal texture and inflection.

One practical note — 30 seconds is the ceiling, and audio past that point is not used. A short, clean, quiet recording will always beat a long, noisy one.

2. Emotional Depth and Natural Prosody

The "robotic" sound of traditional TTS is gone. The model reads context. It naturally incorporates breaths, slight pauses, and the correct emotional tone for the text in front of it. The result is audio that sounds genuinely human, not just human-like.

3. Lightning-Fast Generation

Despite the quality, the current model is remarkably efficient — in our own testing it renders speech several times faster than the audio takes to play, so even long passages come back quickly rather than after a wait.

4. Sixteen Languages, One Voice

A cloned voice is not locked to the language you recorded it in. Qwen-Audio-TTS covers 16 languages, so the same clone can read English, Spanish, Italian, Chinese, Japanese and more without you enrolling it again.

The Premium Experience, Made Simple

We believe state-of-the-art AI shouldn't be locked behind expensive paywalls or complicated coding interfaces.

We've taken the raw power of the Qwen engine and wrapped it in the simpleTTS.ai interface you already know. You get the quality everyone is talking about, without the headache.

Try It Now

The best way to understand the difference is to hear it for yourself.

Log in to your dashboard, upload a short reference clip, and hear what a current-generation clone sounds like.

Related Articles