The Atlas model can take a single image and generate a detailed, photorealistic, minute-long video in 1440p — and you can freeze the frame and generate a new perspective at any time
World Labs calls it a «multimodal autoregressive diffusion transformer,» which is another way of saying that it puts the image or images into an extensive and knowledgeable world model that can «imagine» the parts that are out of the frame and generate an entire world from very little input.
The striking thing about its outputs is how intuitive and realistic they are, bringing single pictures of, say, a cathedral floor together for a walkthrough in the gallery.
It also seems easy to use. Simply upload between one to 12 images, design a path for the camera, and the model does the rest.
This can be useful for walking through works of art for leisure, exploring a path through a landscape of photos, or for robotics training that no longer requires complex, expensive setups to capture a world in a trainable 3D context.
The model will be used to upgrade outputs from their word building Marble app in future iterations, but is not available for public use just yet. It is in «early access with select partners,» and there is an email list to be notified when it becomes available.
Read more: World Labs’ presentation, X launch post. Discussion on Hacker News and r/Singularity.












