MiniMax H3 officially released: unified generation of all modes, 15 seconds of 2K video and 1/3 the price
MiniMax H3 is officially released. It is a universal full-modal generation model that supports unified understanding and output of text/image/video/sound. It can generate up to 15 seconds of 2K video. The price is about 1/3 of the mainstream model.
MiniMax H3 was officially released on July 31st and included in the database on August 2nd. This is a model positioned as "universal full-modal generation" - the unified understanding and output of text, images, videos, and sounds are entered into the same model. The video side supports up to 15 seconds of 2K resolution generation, and the price is about 1/3 of the mainstream model**.
Full-modal unification: reduce the delay and cost of splicing
The architectural direction of H3 is to integrate understanding and generation into a single model, rather than having multiple single-modal models spliced together in the pipeline. The benefits of this "native multi-modal" route are straightforward: it eliminates the delay and cost of conversion between modalities, and is closer to the current "multi-modal + Agent" workload. For MiniMax, this is also a key leap from text/speech single modality to full-modal unified generation.
15 seconds 2K: touching the practical threshold of short video
Video generation length and resolution directly determine usability. The combination of 15 seconds and 2K has reached the threshold for practical use in short video creation, advertising materials and even film and television-level previews, not just demonstration-level "film-making". This means that H3 is not just a demonstration of capabilities in the laboratory, but is heading towards content production scenarios that can be mass-produced.
1/3 Pricing: Who will have the biggest impact?
It provides full-modal capabilities at 1/3 the price of mainstream models, impacting the cost-effectiveness of overseas closed-source multi-modal models. It also directly benchmarks video generation heads such as Keling and Sora in China. For creators, low-priced full-mode means that the threshold for UGC content production has been further lowered - this is also a continuation of MiniMax's consistent "cost-effective approach" approach.
Several directions worth tracking in the future:
- Real Creation Scenario Availability: 15 seconds of 2K’s actual performance in long narrative and character consistency.
- Actual measurement comparison with domestic video models such as Keling: A direct confrontation between video quality and speed.
- Alternative rhythm of all-modal routes: The diversion speed of the unified model to the single-modal scheme.
Reviews