Byte launched a real-time spatial video AI model: under the supervision of Zhang Yiming, it will be released as soon as next month.
2026-09-07 21:51Models🔥 57.3 heat score
4sources
1days unfolding
57.3heat score
9mentions
SummaryAI generated
On September 7, 2026, multiple media sources reported that ByteDance was developing an AI model for real-time spatial video generation, which would be released as soon as next month (late September). The project is overseen by founder Zhang Yiming and is based on Seedance technology. Its goal is to drive the development of the ecosystem and shift the focus of industry competition from hardware parameters to AI models and content platforms. This model supports voice and motion interactions with Pico headsets, capable of generating video streams at 20 frames per second with a delay of 50 milliseconds in the cloud. It allows users to create interactive virtual worlds, similar to Google Genie. The project is intended for live streaming, short videos, and gaming scenarios, by transferring computationally intensive tasks to the cloud to reduce the computing requirements and costs of VR devices. Currently, ByteDance’s spokesperson has not responded officially to this matter.
Zhang Yiming, founder of ByteDance, has joined an elite team composed of world-class scientists dedicated to improving world models. This team includes leading figures in the AI field such as Sam Altman, CEO of OpenAI, and Deborah Harris, chief scientist at DeepMind. Zhang Yiming stated that building a universal artificial intelligence model capable of understanding and predicting the real world is a key direction for the future. This move marks an intensification of competition among global tech giants in basic artificial intelligence research, aiming to accelerate the training and verification of world models through cross-institutional cooperation.
ByteDance plans to release a real-time spatial video generation AI model in late September 2026, developed under the supervision of founder Zhang Yiming. The model is based on existing Seedance technology and is designed for live streaming, short videos, and gaming scenarios. It supports voice and motion interactions with Pico headsets and can generate 3D virtual content at a rate of 20 frames per second with a latency of 50 milliseconds, reducing the computational requirements at the terminal level. This move aims to create an AI-cloud-content-hardware collaborative ecosystem and shift the focus of the VR industry towards competition between models and platforms. The release date has not been finalized yet.
ByteDance is launching an AI model for real-time spatial video generation, with the release expected next month. The model was developed under the direct supervision of Zhang Yiming and is built on Seedance. It supports creating interactive virtual worlds for live streaming, short videos, and games. It can generate videos at a rate of 20 frames per second, with a latency of approximately 0.05 seconds, and can respond to voice or motion commands from Pico headset users. The project aims to move computationally intensive processes to the cloud to reduce VR hardware costs and to explore competition with Meta and Google’s parent company Alphabet in the fields of robotics and autonomous systems. If successful, this model could open up new opportunities for ByteDance in the spatial computing market alongside Meta and Apple. As of press time, a ByteDance spokesperson had not commented.
ByteDance is developing an AI model for real-time spatial video generation, with the expected release next month. Founder Zhang Yiming personally oversees the development, coordinating various departments and AI resources to advance the project. The model is built on Seedance and aims to drive the “flywheel” of the ecosystem, allowing users to create interactive virtual worlds, similar to Google Genie. The model can respond to voice or gesture commands from the Pico headset, generating video streams at 20 frames per second with a latency of 50 milliseconds, supporting spatial computing and human-computer interaction. The architecture will transfer generation tasks to the cloud, reducing the computational requirements and usage costs of VR devices, and shifting the focus of industry competition from hardware parameters to AI models, computing power, and content distribution platforms.