ElevenLabs voice ecosystem is further advanced: Scribe v2 and ElevenAgents Expressive Mode upgrade "speaking" to "expressing"

ElevenLabs continues to expand the voice ecosystem. Scribe v2 speech recognition and ElevenAgents’ expression capabilities have been enhanced. Together with Music v2 and Dubbing v2, the three-layer product structure (creation/conversation/API) has been simultaneously strengthened.

ElevenLabs further advances its voice ecosystem: Scribe v2 and ElevenAgents Expressive Mode, upgrading "speaking" to "expressing"

ElevenLabs will continue to advance its voice ecosystem in 2026. The latest update focuses on two main lines: Scribe v2 speech recognition (STT) and ElevenAgents’ expressive mode enhancement (Expressive Mode), and is superimposed on creative milestones such as Music v2 and Dubbing v2. This company, which positions itself as "AI Communication Platform", is advancing speech from "generation" to "understanding and conversation" through the three-layer structure of ElevenCreative, ElevenAgents, and ElevenAPI.

  • Scribe v2 (STT): Released on 2026-01, officially disclosed that it has reached 98% accuracy on multiple speech recognition benchmarks, making up for the platform's shortcomings on the "speech understanding" side.
  • Expressive Mode for Agents: Launched in 2026-02, allowing customer service/business voice agents not only to "answer correctly", but also to express tone and emotion, improving the naturalness of human-machine conversations.
  • Double milestones on the creative side: Music v2 and Dubbing v2 will be launched in 2026-05, upgrading music generation and multi-language dubbing capabilities as a whole.
  • Three-tier product structure formed: ElevenCreative (creation), ElevenAgents (session), and ElevenAPI (developer) share the same underlying model and asset system.

Version background

The core asset of ElevenLabs is the speech basic model, covering capabilities such as Text to Speech, Speech to Text, Voice Cloning, Dubbing, Music Generation, SFX and Image & Video. Its speech model lineage includes: Eleven Flash (75ms ultra-low latency), Eleven Multilingual v2 (high fidelity), Eleven v3 (high expressiveness) and Scribe v2 (STT, 98% accuracy). The official website annotation supports 70+ languages ​​and 5000+ preset voice libraries.

The main line of the update in 2026 is to put "speech generation" and "speech understanding" into the same product link to reduce the inconsistency in timbre and data island problems caused by the splicing of multiple suppliers in enterprises - this is also a key step in its upgrade from "TTS tool" to "voice infrastructure".

Highlights of this version

Speech understanding side: Scribe v2

  • High Accuracy Transcription: Officially disclosed multiple benchmarks with 98% accuracy, covering multi-language scenarios.
  • Supports closed-loop conversation: STT and TTS collaborate to allow the voice agent to have a complete closed-loop of "understanding and answering" instead of just one-way broadcasting.

Session side: ElevenAgents Expressive Mode

  • Emotional expression ability: Agent can adjust tone and emotional expression according to the context, significantly improving the listening experience in customer service scenarios.
  • Business configuration capability: Provide configurable and monitorable voice and text dialogue agents for customer service and business processes.

Creative side: Music v2 and Dubbing v2

  • Music v2: The music generation capability has been upgraded as a whole, and the creator’s audio asset link has been completed.
  • Dubbing v2: Multi-language dubbing capabilities are enhanced to support the global distribution of film, television and content.

Meaning for developers

For domestic developers and product teams, there are two points worth paying attention to in the ecological evolution of ElevenLabs. First, voice Agent is becoming a new carrier for customer service, outbound calls, and marketing, and "expressiveness" is a key variable that distinguishes it from traditional IVR - capabilities such as Expressive Mode directly affect users' acceptance of AI customer service. Secondly, the integrated solution of STT + TTS + Agent has more advantages than "multiple assembly" in terms of timbre consistency and data management, but we must also pay attention to the compliance and privacy requirements of voice data.

In comparison, Whisper on the open source side provides another path to free self-hosting on STT, while the value of ElevenLabs lies in packaging identification, generation and sessions into manageable enterprise-level services. Both are suitable for "self-built data pipelines" and "quick launch" teams respectively.

Use path suggestions

  • Creator: Start with ElevenCreative to experience voice cloning, dubbing and music generation, and verify your content production workflow.
  • Enterprise Customer Service/Outbound Call: Evaluate the Expressive Mode and monitoring capabilities of ElevenAgents, and expand the capacity after piloting with small traffic.
  • Developer: Integrate TTS/STT/Agent capabilities through ElevenAPI, and focus on the cost model of seven levels of subscriptions (Free to Enterprise) and usage Credits.

Directions worth tracking

  1. Chinese and dialect coverage of Scribe v2: Measured performance of multi-language transcription quality in domestic localization scenarios.
  2. Expressive Mode implementation scenarios: Adoption rate and user acceptance data in customer service, audio content and other scenarios.
  3. Enterprise Compliance and Data Security: The compliance path for voice data used in China is a key prerequisite for determining whether the domestic team can access it.
Copyright: Content sourced from ElevenLabs official release . This platform has compiled and organized this content for informational purposes and learning exchange only. If there are any copyright concerns, please contact us for resolution.

Reviews

  • Loading reviews...