The research team released WorldSculpt, aiming to use single-object 3D generation to generate messy scenes containing hundreds of objects. This method adapts Pixal3D models to multi-view observations, and generates objects in multiple poses through multi-view conditional paths, allowing it to generalize to large scenes with severe occlusion without scene-level training. The team established the UE-MeshyScene benchmark dataset, which includes real-photo scenes, object-level annotations, and true grids. Experiments show that WorldSculpt outperforms existing methods in single-object, multi-object control, and this benchmark test, and its advantages become more evident as the complexity of the scene and occlusion increase. Additionally, this method supports converting 3DGS worlds generated by Marble, HY-World 2.0, etc., into composite grid scenes.
WorldSculpt releases models and establishes the UE-MeshyScene benchmark dataset; experimental results are pending verification.
Coverage · reports per dayLANGUAGE SPLIT
Entity relations
Integrated timelineUNIFIED TIMELINE
2026-09-04
The research proposal for WorldSculpt begins to take shape.
The paper on WorldSculpt is first published on arXiv, proposing dense scene generation without scene-level training using multi-view conditional paths.
2026-09-07
Publication of the WorldSculpt paper and disclosure of details
Hugging Face Papers and arXiv respectively publish the WorldSculpt paper, introducing its method for generating dense occluded scenes using single-object priors and Pixal3D adaptation technology.
The research team proposed WorldSculpt, which achieves scalable combination mesh reconstruction of dense scenes with hundreds of objects by adapting 3D generation priors to multi-view observations for single objects. This method utilizes Pixal3D models and introduces multi-view conditional paths, anchoring object generation in multiple pose observations; although the model is only fine-tuned on single objects in standard space, it can generalize to heavily occluded large scenes without scene-level training. The research team also released the UE-MeshyScene benchmark dataset, which includes real-photo scenes with hundreds of objects, along with per-object annotations and true meshes. Experiments show that this method outperforms existing methods in single-object, multi-object control, and UE-MeshyScene evaluations, and its advantages become more apparent as scene complexity and occlusion increase. Additionally, the study demonstrates the broad applicability of converting generated 3DGS worlds (such as Marble and HY-World 2.0) into combined mesh scenes.