SpatialCrafter:使用生成 3D 代理进行单图像世界建模
SpatialCrafter: Single Image World Modeling with Generative 3D Proxies
摘要
Explorable image-to-scene generation is essential for applications in gaming, robotics, and virtual reality.基于视频扩散模型 (VDM) 的现有方法通常依赖于不完整的条件信号,例如稀疏点云或 2D 全景图,导致随机幻觉、长期漂移和次优 3D 一致性。我们提出了 SpatialCrafter,这是一种新颖的两阶段框架,通过引入用于高保真图像到场景生成的全局 3D 代理来解决这些问题。 Specifically, we decompose the generation process into global proxy generation and appearance refinement.对于代理生成,我们提出了一个点锚定稀疏结构~(PaSS)流模块,该模块可预测空间对齐且几何一致的 3D 代理。为了进行外观细化,我们将 VDM 重新构建为生成延迟细化器,它根据代理定义的场景几何图形合成高频真实感细节。为了更好地将代理与预训练的 VDM 集成,我们引入了并行几何注入和代理感知损坏训练策略,这些策略在不破坏预训练生成流形的情况下提高了代理工件的鲁棒性。 Furthermore, as no suitable dataset exists for this explorable scene generation task, we construct a new large-scale dataset of 115K scenes. To the best of our knowledge, it is the first hybrid dataset for image-to-scene generation.对合成数据集和真实世界数据集的大量实验表明,SpatialCrafter 的性能优于最先进的方法,减轻了长期漂移,并在快速相机运动和极端视点变化下保持鲁棒性和一致性。 Our project page: \href{https://fangchuan.github.io/SpatialCrafter/}{fangchuan.github.io/SpatialCrafter/}