We’re open-sourcing HY-World 2.0, a multimodal world model that generates, reconstructs, and simulates interactive *3D worlds* from text, images, and videos. Outputs can be integrated into game engines and embodied simulation pipelines. Key highlights: 🔹 One-click world generation Turn text or image into interactive 3D worlds automatically. 🔹 Pipeline-ready 3D outputs Editable 3D worlds for Unity and Unreal Engine, with standard 3D exports including mesh, 3DGS, and point clouds. 🔹 Unified world model system One model family for world generation and reconstruction across synthetic and real-world scenes. 🔹 Interactive character mode Explore generated 3D worlds in real time with physics-aware movement and collision support. ✨ Apply for access: https://t.co/swscD5KGu2 🔗 GitHub: https://t.co/XpUKjBtK5n 🤗 Hugging Face: https://t.co/tv8hOPYABj 📄 Technical Report:
Technical highlights from HY-World 2.0 👇 - 3D-first world modeling: a unified framework for world generation and reconstruction, built around spatial understanding in 3D. - HY-Pano 2.0: scales panorama generation for high-fidelity 360° world initialization from single images — no camera metadata required. - WorldNav (spatial agent + navmesh): semantic-aware trajectory planning that combines VLMs with navmesh for coherent, collision-aware exploration. - WorldStereo 2.0: keyframe-based world expansion in a latent space with spatially consistent memory for stable novel view synthesis. - WorldMirror 2.0: unified 3D reconstruction that composes multi-view predictions into accurate, navigable 3DGS assets. - WorldLens: a high-performance, engine-agnostic 3DGS renderer for interactive exploration with lighting and collision support. 📄 Read the full technical report:

Unlike video-based world models, the output is editable 3D geometry that loads into a game engine or an embodied-simulation pipeline, with weights, code and a technical report published alongside.
Checking sign-in…
Loading comments…