Li Feifei’s first multimodal world model, Atlas, was officially released: Pre-trained from scratch, and it can build a world with just a few images.
The team of computer scientists at Stanford University, led by Li Feifei, officially released the first multimodal world model, Atlas. This model uses pre-training from scratch, and only a few images are required to create a complete view of the world. Atlas can understand objects, scenes, and spatial relationships in images, and it also supports the generation of new image content. This achievement marks a major breakthrough in the field of deep integration of artificial intelligence in vision and language, laying an important foundation for the subsequent development of general artificial intelligence systems.