Introducing Atlas: The worlds first multimodal world model that generates image and video frames with pixel-perfect camera control and reconstructs them in 3D. Model the world, move the camera, and simulate space & time.
We pre-trained Atlas from scratch to take multimodal inputs, including camera movement, and turn it into 3D grounded views. Atlas puts you in control: direct the views, reconstruct real spaces with your inputs, and build explorable worlds. Read more: https://www.worldlabs.ai/blog/atlas
By positioning multiple input views within the models spatial context, Atlas allows you to generate image and video frames with precise control. Here, we hand-designed a camera trajectory to generate a 1 minute video at 1440p resolution from seven reference images.
Atlas also outputs explicit 3D from one to many input images, beating top open source reconstruction models. Passing more images gives Atlas more context: the more it sees, the less it imagines.
Atlas turns a handful of ordinary photos into controllable 3D with depth, removing the specialized scanning hardware step from building simulation environments and multiview capture setups.
Checking sign-in…
Loading comments…