Miss Jia’s conversation with Li Zhifei: Understanding Sora, reproducing Sora
2026-09-04 08:00Models🔥 42.2 heat score
1sources
1days unfolding
42.2heat score
5mentions
SummaryAI generated
With the release of the video generation model Sora by OpenAI, domestic AI practitioners’ attitude has shifted from skepticism to attempts at replicating it. Li Zhifei, founder and CEO of Mengchuangwen, compiled a complete technical architecture after studying 32 papers and technical reports. He believes that Sora does not represent a breakthrough in principle; the core challenge lies in the duration, quality, and consistency of generated videos. This confirms the applicability of Transformer and the law of scale in the video domain, suggesting that Sora is in a transitional stage between GPT-2 and GPT-3. Li Zhifei points out that since OpenAI has not disclosed details such as the specific implementation of its encoding/decoding mechanisms and the mechanism for generating long videos, domestic companies find it difficult to replicate Sora, and they face even more severe challenges related to video copyright compliance compared to ChatGPT. Currently, OpenAI’s team has clearly stated that they will not release Sora soon, and its actual capabilities may be superior to the existing demonstration version.
Domestic AI practitioners are shifting from skepticism to attempts at replicating the video generation model Sora released by OpenAI. Li Zhifei, founder and CEO of Menwen, has independently studied 32 related papers and one technical report provided by OpenAI, and has pieced together a complete technical architecture diagram of Sora. Li believes that Sora does not represent a breakthrough in principle; its core impact lies in the duration, high quality, and consistency of generated videos. He speculates that Sora can be seen as a transitional stage between GPT-2 and GPT-3, demonstrating the applicability of Transformer and the law of scale in the video domain. Regarding the difficulty for domestic companies to catch up, Li points out that Sora is difficult to replicate due to undisclosed technical details (such as the specific implementation of codecs and the mechanism for generating 60-second videos), and it faces more severe challenges related to video copyright compliance than ChatGPT. Currently, OpenAI’s team has clearly stated that they will not release Sora soon, and its true capabilities may be higher than those shown in existing demos.