MiniMax Speech 2.8 and Music 3.0: Dual-line upgrade of speech and music generation

MiniMax released the Speech 2.8 voice model and Music 3.0 music model in June, both providing APIs through platform.minimax.io; together with M3 and H3, they form a multi-modal matrix, showing its dual-line commercial layout in C-side entertainment and B-side APIs.

MiniMax will release the Speech 2.8 speech model and Music 3.0 music model in June 2026, further strengthening its capabilities in the fields of speech and music generation. Both provide API and experience access through the platform.minimax.io platform, and together with M3 and H3, form MiniMax's multi-modal basic model matrix.

Two models, two business lines

Speech 2.8 is oriented to speech synthesis and cloning scenarios, and Music 3.0 is oriented to music generation scenarios. Although they share the technical base of "audio generation", they correspond to two completely different commercialization paths of MiniMax:

  • B-side API: The capabilities of voice cloning and music generation are output to developers and corporate customers through the API - it can be used in dubbing, audio content, music production and other scenarios.
  • C-end entertainment: Cooperate with MiniMax Audio and Talkie (AI companion product) to package voice and music capabilities into entertainment products for C-end users.

One line sells capabilities and the other line sells experience. This is MiniMax's typical approach of "making both infrastructure and consumer products".

Another link in the multimodal matrix

Putting Speech 2.8/Music 3.0 back into MiniMax's overall landscape: The official positioning of the series of models is "a universal basic model with powerful code and agent capabilities, ultra-long context processing capabilities, and the ability to understand, generate and integrate multi-modal modalities such as text, audio, images, videos, and music." Voice and music are not isolated functions, but the carrier of the "audio" dimension in this multi-modal matrix - when M3 handles long tasks, H3 generates videos, Speech 2.8 synthesizes speech, and Music 3.0 generates music, MiniMax's full-modal puzzle is taking shape.

From an industry perspective, competition in the voice and music generation tracks is equally fierce—overseas players such as ElevenLabs and Suno have already taken a strong lead. MiniMax's differentiation lies in "full-modal integration": developers can call text, voice, music and video capabilities on the same platform and the same API system, reducing the technical debt of multi-vendor assembly. For domestic audio and music creators, the MiniMax API also provides a domestic alternative.

Several directions worth tracking in the future:

  1. Compliance Boundaries of Voice Cloning: Authorization and review mechanism for cloning capabilities under deep forgery governance.
  2. Music 3.0 Copyright Policy: How to define the copyright ownership and commercial authorization of generated music.
  3. Talkie’s C-side performance: AI companion products’ commercial support for voice capabilities.
  4. Full-modal API integration: The experience and cost of calling multi-modal capabilities on the same platform.
Copyright: Content sourced from MiniMax official . This platform has compiled and organized this content for informational purposes and learning exchange only. If there are any copyright concerns, please contact us for resolution.

Reviews

  • Loading reviews...