Back to News

World Labs unveils Atlas, a multimodal world model with pixel-level camera control and 3D reconstruction

#world-models#multimodal#3d-reconstruction#camera-control

World Labs, founded by Fei-Fei Li, unveiled Atlas, described as the world's first multimodal world model. Atlas generates images and video frames with pixel-perfect camera control, performs 3D reconstruction from one to dozens of images, and supports spatiotemporal simulation for applications such as robotics and visual effects. The model is a multimodal autoregressive diffusion transformer, capable of outputting up to 1-minute 1440p video and generating 360-degree panoramas from text.

Coverage timeline

  1. Techmeme

    World Labs : Fei-Fei Li's World Labs unveils Atlas, a multimodal world model that generates image/video frames with pixel-perfect camera control and reconstructs them in 3D — World models generate, reconstruct, and simulate any possible world. They understand how worlds appear, behave …

  2. 量子位量子位

    李飞飞发布:全球首个多模态世界模型 鱼羊 2026-09-02 09:07:13 来源: 量子位 一张图补全3D世界,还能给机器人造训练场 鱼羊 发自 凹非寺 量子位 | 公众号 QbitAI 「全球首个」 多模态世界模型 ,来了! 来自李飞飞World Labs。 World Labs对这个名为 Atlas 的新一代世界模型,用词可谓毫不克制,「全球首个」之外,slogan也明确强调出,这一次,可不只是生成一段可互动视频了哟: 建模世界,移动相机,模拟空间和时间。 就是说,Atlas能够以像素级的相机控制生成图像和视频帧,并完成3D重建。并且 空间 和 时间 都能理解。 李飞飞本人将其视作World Labs的「里程碑式」成果,直接一手置顶: Atlas为 从视觉特效到机器人 的众多应用场景打开了大门。 飞飞高徒、英伟达机器人主管Jim Fan也来捧场,指出:这是机器人领域Real-to-Sim的一大步。 具体来说,Atlas是一个 多模态自回归扩散Transformer ,它的能力包括: 相机控制生成 :仅需一张图片,Atlas即可生成具有像素级相机控制的图片和视频,最高可输出1分钟、1440p的视频。 空间重建 :Atlas可以基于一张到数十张输入图像,重建真实世界场景。既可以生成新视角下的图像帧,也能输出显式3D表示,其效果超过当前专门针对3D重建训练的最先进模型。 时空模拟 :Atlas可以根据输入视频同时建模空间和时间,可以转换已有视频中的观察视角,制作具有戏剧性的视觉效果,同时支持机器人Real-to-Sim工作流。 图像生成 :Atlas可以根据文本生成图片和360°全景图,能够遵循复杂Prompt、准确渲染文字,并生成大量不同的视觉风格。 对于 机器人模拟 ,仅需喂Atlas几张照片,它就能生成逼真的RGB和深度数据。这样一来,机器人就能在更多不同的模拟空间中进行训练和测试。 从头训练的多模态世界模型 技术细节上,Atlas是一个 全能模型 (omni model)。 它的目标,是在一个统一的模型架构中,同时处理多任务、多种输入和输出数据。 与此同时, 空间控制是整个模型的核心 。 为了实现这些目标,李飞飞团队设计了一种全新的基础架构,来作为未来世界模型的基础—— 多模态自回归扩散 Transformer 。 文本、图像、视频、3D数据……各种输入