Technology
World Labs announces Atlas, a multimodal world model for spatial intelligence
The autoregressive diffusion transformer generates 3D scenes, simulations, and controlled video from multimodal inputs.
The short version
- World Labs unveiled Atlas, a base world model pretrained to natively process text, images, video, and 3D data within a shared spatial context.
- The system targets developers, visual effects creators, and roboticists by generating up to 1440p camera-controlled video, 3D Gaussian splats, and Real-to-Sim environments.
- Atlas is designed to power future iterations of World Labs products, including Marble, though commercial availability and deployment timelines were not detailed.
Key facts
- World Labs introduced Atlas as an omni world model pretrained from scratch using a multimodal autoregressive diffusion transformer architecture.[Hacker News]
- The model accepts native camera poses alongside text, images, and depth maps, enabling video generation up to 1 minute at 1440p resolution with explicit trajectory control.[Hacker News]
- Atlas outputs explicit 3D representations, such as point clouds and 3D Gaussian splats, to reconstruct real-world scenes from sparse single or multiple reference images.[Hacker News]
- The architecture combines rectified flow diffusion denoising with transformer-based sequence modeling to simulate space-time dynamics for visual effects and robotics Real-to-Sim workflows.[Hacker News]
- World Labs confirmed that Atlas will serve as the underlying foundation for future versions of Marble and related company tools.[Hacker News]
What remains uncertain
- Public access dates, API availability, compute requirements for end users, and pricing for Atlas have not been specified.[Hacker News]
- World Labs claims Atlas outperforms specialized state-of-the-art 3D reconstruction models, but external benchmark evaluations were not provided.[Hacker News]