ElevenLabs releases Dubbing v2: AI dubbing hits the main battlefield of "film and television localization"

ElevenLabs released Dubbing v2 on May 28, which has been fully upgraded in terms of speech naturalness, timbre fidelity, lip synchronization accuracy, emotional consistency and long content stability. Together with Eleven v3 and Scribe v2, it forms a complete "recognition-translation-synthesis-synchronization" dubbing pipeline, intensifying the global competition for multi-language content.

Automatically translate a video into dozens of languages ​​while retaining the timbre, mouth shape and emotion of the original speaker - this is what ElevenLabs's Dubbing product has been doing, and it is also a core requirement for film and television localization and content overseas. Dubbing v2, released on May 28, takes this capability to the next level.

What is upgraded in v2?

The upgrade direction of Dubbing v2 covers almost every aspect of the dubbing pipeline: voice naturalness, timbre fidelity, lip sync accuracy, emotional consistency, and long content stability, and strengthens multi-language/multi-accent support. For the content overseas team, the stability of long-form content is particularly critical - for a 90-minute movie, if the dubbing goes off track at the 80th minute, all previous efforts will be wasted.

A complete dubbing pipeline

Dubbing v2 is not an isolated capability. It collaborates with Eleven v3 speech model, Scribe v2 (ASR) and other technologies to form a complete dubbing pipeline of "recognition-translation-synthesis-synchronization". The significance of this pipeline is that it is not "capable of dubbing", but "capable of industrialized batch dubbing" - this is exactly what is needed in scenarios such as film and television, education, corporate content, and overseas gaming.

From an industry perspective, AI dubbing/localization is an important pillar of ElevenLabs' commercialization, and competition in this track is intensifying: HeyGen, Synthesia, etc. are also making efforts in the video translation/lip synchronization track. The release of Dubbing v2 has raised the threshold of this track yet another level - when "tone fidelity + lip synchronization + long content stability" becomes standard, the competition will shift from "can it be matched?" to "whether it is matched well, whether it is fast or not, and whether it is expensive or not." For domestic AI dubbing and video exporting tools (Alibaba and Byte's dubbing products), this is both direct pressure and a direction reminder: "Multi-language content globalization" is the core application scenario of AI audio and video in 2026, and whoever can solidify this pipeline will hold the key to exporting content overseas.

Several directions worth tracking in the future:

  1. Actual accuracy of lip synchronization: The naturalness of lip synchronization in multiple languages.
  2. Long content stability: Failure rate of long video dubbing and the need for manual intervention.
  3. Pricing and localization costs: Can the unit cost of batch dubbing support large-scale overseas expansion?
  4. Response to domestic video export tools: How do Alibaba and Byte's dubbing products follow up on the "complete pipeline" capabilities.
Copyright: Content sourced from ElevenLabs official blog . This platform has compiled and organized this content for informational purposes and learning exchange only. If there are any copyright concerns, please contact us for resolution.

Reviews

  • Loading reviews...