Eleven v3 is officially available: AI voice enters a new stage of "difficult to distinguish real people"
ElevenLabs announced on February 2 that its flagship speech model Eleven v3 is officially available (GA), reaching a new level in speech naturalness, emotional expression, multi-language capabilities and stability. It forms a complete audio technology stack with Scribe v2, Dubbing v2, and Music v2, and is the technical base for "the most realistic AI voice".
On February 2, ElevenLabs announced that its flagship speech model Eleven v3 is generally available. For a company whose mission is "realism", v3's GA means that its speech synthesis technology has entered a new generation - speech naturalness, emotional expression, multi-lingual capabilities and stability have reached new levels at the same time, becoming the technical base for "the most realistic AI voice".
Technical context from v1/v2 to v3
ElevenLabs’ speech model iteration path is clear: v1/v2 lays the foundation for speech cloning and TTS, while v3 strengthens long-text consistency, emotion/tone control, and multi-language (covering dozens of languages) capabilities. At the same time as v3, ElevenLabs also released/iterated Scribe v2 (speech to text, January 2026), Dubbing v2 (dubbing, May 2026), Music v2 (May 2026), etc., forming a complete audio technology stack of "speech synthesis + recognition + translation + music".
The foundation of a business empire
The significance of v3 is not just about technology. Combining its Series D financing (valued at US$11 billion) and US$500 million in ARR, ElevenLabs is building a voice AI business empire based on v3: voice agent, dubbing, audio books, customer service - these scenarios are all based on the premise that "AI voice is lifelike enough". The GA of v3 is equivalent to putting a "qualified" label on the foundation of these business scenarios.
From an industry perspective, Eleven v3's GA marks that AI voice quality has entered a new stage of "hard to distinguish from real people". This is not only a technical victory, but also brings new questions: When the AI voice is no different from a real person, how to establish "voice trust"? How can the governance of deepfakes keep up? For domestic voice AI companies (Byte, Alibaba, Mobvoi, etc.), v3 is both a competitive reference - fidelity has reached a new level - and a safety warning - the other side of fidelity is the risk of abuse. In the second half of the AI voice track, the competition will not only be "whether it looks like it", but also "whether it can be used responsibly".
Several directions worth tracking in the future:
- The actual listening difference between v3 and v2: Comparison of naturalness and emotional expression in real scenes.
- Multi-language capability coverage: Among dozens of languages, the sound quality level of Chinese and other languages.
- Deep Forgery Governance: ElevenLabs’ voice watermarking and traceability solution.
- Follow-up of domestic TTS: Can local voice manufacturers narrow the gap in naturalness.
Reviews