Sora’s road to productization: from research preview to ChatGPT video capabilities and its unsolved physics challenges

Sora is OpenAI's text generation video model, which can generate high-quality videos up to one minute long. From the research preview in February 2024 to its inclusion in the ChatGPT product system, Sora's diffusion model architecture and unresolved physical simulation limitations are worth sorting out.

Sora is OpenAI's text-generated video model that generates videos up to a minute long while maintaining visual quality and following user prompts. As one of the most influential research milestones in the field of video generation, the path it has taken—from a research preview in February 2024 to now becoming a part of the ChatGPT product system—is itself a microcosm of the AI ​​video industry.

Technical Base: Diffusion Model + Video "Patch"

Sora's technical path can be summarized into two layers. First, it is essentially a diffusion model: first generate a video similar to static noise, and then gradually transform it into a coherent picture through multi-step iterative denoising. Second, it uses the Transformer architecture to represent videos and images as a collection of "patch" - a token similar to GPT. This brings several useful capabilities:

  • Generate the entire video at once, or expand and generate existing videos;
  • Animate and add frames to existing pictures/videos;
  • Multi-shot generation and multi-character scene understanding.

Productization: From Research to Part of ChatGPT

Now Sora has been included in the ChatGPT product system (sora.com points to sora.chatgpt.com) and has become an important part of the OpenAI multi-modal matrix (GPT-5.6, Codex, Sora, Whisper). This means that it is no longer a research preview only for red teams and artists to test, but a product capability accessible to ordinary users.

Unsolved physics problems

As a review, it is also necessary to point out the practical limitations of Sora: it still has obvious shortcomings in complex physical simulation, causality (such as bitten biscuits should leave bite marks) and spatial details (such as left and right directions). These are typical manifestations of the video generation model's "image but not right enough" - the picture is beautiful, but the operating rules of the world are not really understood by the model.

From an industry perspective, the value of Sora is not only the product itself, but the technical route it defines for "Wensheng Video": while later Seedance, Kling, and Veo all surpassed them in their respective fields, Sora's pioneering still lies in its earliest proof of the feasibility of the "diffusion + Transformer + patch" path. For domestic researchers, the evolution of Sora also prompts a judgment standard - the next milestone in video generation is most likely not longer duration, but the correct modeling of physical laws.

Several directions worth tracking in the future:

  1. Breakthrough in physical simulation capabilities: Can subsequent versions of Sora solve the problems of causality and spatial consistency.
  2. Depth of integration in the ChatGPT system: Sora’s linkage scenarios with GPT-5.6 and Codex.
  3. The increase in the one-minute upper limit: Will long video generation become the focus of the next round of capability competition?
  4. Gap with Seedance, Kling, and Veo: Will OpenAI strive to catch up in video generation capabilities?
Copyright: Content sourced from OpenAI official . This platform has compiled and organized this content for informational purposes and learning exchange only. If there are any copyright concerns, please contact us for resolution.

Reviews

  • Loading reviews...