🌍 ElevenLabs Dubbing v2: Key Takeaways
• Automatic Dubbing in 90+ Languages: A significant expansion from the ~70 languages supported in v1.
• Performance Preservation: Retains the original tone, emotion, cadence, and delivery across target languages.
• Audio-to-Audio Model: Eliminates the need for intermediary transcripts, conditioning directly on the source audio.
Hello from the Sonetho team! ⚡
On May 29, 2026, ElevenLabs officially announced the GA (General Availability) release of our next-gen dubbing model, Dubbing v2.
Following our Music v2 announcement on May 27, this is our second major launch of the week—and it represents far more than just an increase in language support.
The engine under the hood has been completely redesigned.
v1 vs. v2 — What’s changed?
Compare this to our v1 guide from a few months ago, and you’ll see the shift immediately.
Feature | Previous (v1) | New (v2) |
|---|---|---|
Model Architecture | Transcript-based | Direct Audio-to-Audio |
Emotion & Nuance | Limited | Original Speaker Fidelity |
Video Sync | Required manual intervention | Sync-aware automation |
Voice Cloning | Required separate PVC training | Built-in Auto-Cloning |
Languages | 70 | 90+ |
As the table shows, v2 outperforms in every category — now it's time to test it yourself.
Experience automatic voice cloning, 90+ languages, and emotional fidelity with the Creator plan at a discounted $11/mo.
Try Dubbing v2 free — no credit card →
5 Key Pillars of the Official Release
1. Performance-aware Dubbing — Preserving the acting
Official messaging: "Transfer the speaker’s tone, emotion, pacing, delivery, and intent seamlessly into another language."
With v1, dubbing could sometimes sound like a "mechanical read."
In v2, the emotional nuances—think of a sigh, a chuckle, or a dramatic pause—are preserved with incredible accuracy. It’s a game-changer for creators looking to reach global audiences.
2. Audio-to-Audio Model — Removing the transcript step
Official messaging: "We no longer rely on transcripts; we condition directly on the source audio."
This is the technical heart of v2. It handles background noise and music much better, and, crucially, it excels at speaker separation. If your video features overlapping speech, v2 maintains distinct speaker identities.
3. Sync-aware Translation — The translated audio’s start and end times are automatically adjusted to fit the visual content. This is vital when moving from English to languages with different syllable counts, ensuring the dubbed audio feels natural rather than rushed or laggy.
4. Built-in Automatic Voice Cloning — You no longer need to manually train a voice clone to match the speaker. It’s now seamlessly integrated into the dubbing process.
We’ll be running our own benchmarks to see how it compares to our dedicated PVC (Professional Voice Cloning).
5. 90+ Languages — A significant leap from v1. Whether you are targeting Spanish, French, or Japanese, the quality consistency across non-English language pairs has seen a massive improvement.
7-Day Special — 30 minutes of free dubbing for paid plans
To celebrate the launch, we’re adding 30 minutes of free dubbing credit to every paid plan.
It applies automatically—but this is only for the first 7 days after launch.
⏳ The 7-day bonus is applied automatically at checkout — no coupon code required.
New signups get the Creator plan for $22 → $11 (50% off), plus the 30-minute bonus. Get both by signing up today.
Start free with your 7-day dubbing bonus →
💡 Aside from the 7-day bonus, here is how you can always secure a 50% discount as a new user → How to get 50% off the ElevenLabs Creator plan in 2026
Plan | Base Dubbing | + 7-Day Bonus | Total |
|---|---|---|---|
Free | 10 min (Trial) | — | 10 min |
Starter ($5) | 30 min | + 30 min | 60 min |
Creator ($22, Sale $11) | 2 hours | + 30 min | 2.5 hours ← Recommended |
Pro ($99) | 10 hours | + 30 min | 10.5 hours |
Coming Up Next
We’ll be deep-diving into these five pillars in our next post. In the meantime, if you want to hear it for yourself, check out the official samples on the ElevenLabs site to get a feel for the v2 quality!
🎧 Listen to the official Dubbing v2 demos
📚 Recommended Reading
The Ultimate Guide to ElevenLabs Dubbing: Translate & Dub to 90+ Languages
ElevenLabs Tips · v1 Workflow Rundown
99% Sync Accuracy: The Secrets of Animation Dubbing (Clip vs. Track vs. IVC)
ElevenLabs Tips · Using Auto-Cloning in v2
"Is that AI?" Animation Dubbing from Your Home Studio
ElevenLabs Tips · Real-world Dubbing
🚀 Final Thoughts
If you're serious about taking your content global, now is the perfect time to experiment. The "robotic" barrier that existed in earlier iterations is rapidly dissolving—v2 is a massive step toward true, natural-sounding localization.
🌎 Start Dubbing v2 free — no credit card →
※ The links above are official affiliate links for ElevenLabs.
❓ Frequently Asked Questions (FAQ)
Q1. What has changed in Dubbing v2?
A. The underlying model architecture has been completely overhauled.
While v1 was transcript-based, v2 uses an Audio-to-Audio method that conditions directly on the source audio.
The number of supported languages has increased from around 70 to over 90.
Video syncing has evolved from manual post-editing to automatic alignment, and voice cloning is now built-in rather than requiring a separate PVC training process.
Q2. What does it mean that the original performance is preserved?
A. The official stance is that the speaker's tone, emotion, pacing, delivery, and intent are carried over into the target language.
In v1, dubbing from Korean to English often sounded like a mechanical reading.
Whether non-verbal cues like sighs or laughter are fully captured is something you should verify for yourself.
Q3. What is Sync-aware Translation?
A. This is a feature that automatically aligns the start and end of the translated audio with the original video.
The key factor is how well it matches lip movements or scene transitions in languages where syllable counts differ significantly, such as Korean and English.
Q4. Do I need to create a separate voice clone?
A. It has been announced that v2 automatically generates translated audio in the original speaker's voice without any additional setup.
You may want to verify the quality gap between this and traditional PVC cloning separately.
Q5. How do I receive the 30 minutes of free dubbing for 7 days?
A. To celebrate the launch, 30 minutes of bonus dubbing time is added to all paid plans.
It is applied automatically without the need for a separate application or coupon code.
Please note that this is only available for 7 days starting from the launch date.
Q6. What are the dubbing limits for each plan?
A. The Free plan provides a 10-minute trial with no bonus included.
The Starter ($5) plan includes a base of 30 minutes plus a 30-minute bonus, totaling 60 minutes.
The Creator ($22, discounted to $11) plan includes a base of 2 hours plus a 30-minute bonus, totaling 2 hours and 30 minutes.
The Pro ($99) plan includes a base of 10 hours plus a 30-minute bonus.
Happy creating!
The Sonetho Team ⚡