Today we are sharing three new research papers, each exploring a new way to generate 3D content by leveraging large-scale generative models and 2D priors. These projects were led by our incredible interns @HaoZhang623 @BDuisterhof @DrTunnels [1/4]
World Tracing predicts full 3D from a single image. It outputs a stack of depth values for each input pixel, peeling the world into layers and predicting them with a diffusion model. This predicts full 3D (even occluded surfaces) while remaining faithful to the image. [2/4] https://t.co/sIm3DdOfhr
Modality Forcing adapts text-to-image models to reason jointly about text, images, and depth. It shows that text-to-image is a scalable pretraining objective for 3D reasoning, and how text-to-RGBD, depth estimation, and depth-to-image can be unified in a single model. [3/4] https://t.co/HfYOdgObNw
Flex4DHuman lifts monocular video into dynamic 4D Gaussians. A video diffusion model is finetuned to generate synchronized multiview videos which are distilled into 4D Gaussians. With this method, a video of a person dancing can be lifted to 4D and composted into a 3D world. https://t.co/WUQnfSEq0k
The shared finding is that 2D generative transfers to 3D reasoning: text-to-image checkpoints can be adapted to depth and RGBD, and diffusion can infer occluded geometry faithfully enough to reconstruct a scene from one image.
Checking sign-in…
Loading comments…