AuraTracer智迹闻
中文

EVENT DOSSIER

Paper page - WorldSculpt: Generating Compositional Worlds from Grounded Videos

WorldSculpt releases models and establishes the UE-MeshyScene benchmark dataset; experimental results are pending verification.

2026-09-07 21:57 Models across 2 days 🔥 58.2 heat score hf-papers #20
3sources
2days unfolding
58.2heat score
6mentions
SummaryAI generated

The research team released WorldSculpt, aiming to use single-object 3D generation to generate messy scenes containing hundreds of objects. This method adapts Pixal3D models to multi-view observations, and generates objects in multiple poses through multi-view conditional paths, allowing it to generalize to large scenes with severe occlusion without scene-level training. The team established the UE-MeshyScene benchmark dataset, which includes real-photo scenes, object-level annotations, and true grids. Experiments show that WorldSculpt outperforms existing methods in single-object, multi-object control, and this benchmark test, and its advantages become more evident as the complexity of the scene and occlusion increase. Additionally, this method supports converting 3DGS worlds generated by Marble, HY-World 2.0, etc., into composite grid scenes.

Related eventsRELATED EVENTS
Quick factsQUICK FACTS
Hundreds of objects
Pixal3D
Key entitiesKEY ENTITIES
HY-World 2.0MarblePixal3DUE-MeshySceneWorldSculpttaesiri

Event frameEVENT FRAME

Research

政府 · 科研机构across 2 days

Status

WorldSculpt releases models and establishes the UE-MeshyScene benchmark dataset; experimental results are pending verification.

Coverage · reports per dayLANGUAGE SPLIT

Entity relations
Pixal3D × UE-MeshyScene3Pixal3D × WorldSculpt2UE-MeshyScene × WorldSc…2HY-World 2.0 × Marble2HY-World 2.0 × Pixal3D2HY-World 2.0 × UE-Meshy…2

Integrated timelineUNIFIED TIMELINE

  1. 2026-09-04

    The research proposal for WorldSculpt begins to take shape.

    The paper on WorldSculpt is first published on arXiv, proposing dense scene generation without scene-level training using multi-view conditional paths.

  2. 2026-09-07

    Publication of the WorldSculpt paper and disclosure of details

    Hugging Face Papers and arXiv respectively publish the WorldSculpt paper, introducing its method for generating dense occluded scenes using single-object priors and Pixal3D adaptation technology.

    2 reports

SignalsSIGNALS

Keyword heat
  • Pixal3D3
  • UE-MeshyScene3
  • WorldSculpt2
  • Marble2
  • HY-World 2.02
  • taesiri1

All reports (3)SOURCES

A arXiv cs.CV en 2026-09-05 01:59

WorldSculpt: Generating Compositional Worlds from Grounded Videos

研究者提出 WorldSculpt,旨在生成包含数百个物体的杂乱场景的组成式 3D 表示。该方法通过适配强单物体 3D 生成先验至多视角观测,实例化为 Pixal3D,并引入多视角条件路径将物体生成锚定在多个位姿观测上。尽管模型仅在标准空间对单物体进行微调,却能无需场景级训练泛化至含严重遮挡的大场景。研究团队还建立了 UE-MeshyScene 基准,包含数百物体的密集杂乱场景、逐物体标注及真实网格。在单物体、受控多物体及 UE-MeshyScene 评估中,该方法表现优于先前的方法,且随着场景复杂度和遮挡增加优势更显著。此外,该方法展示了将 Marble、HY-World 2.0 等生成的 3DGS 世界转换为组成式网格场景的广泛适用性。

A arXiv cs.CV en 2026-09-07 12:00

WorldSculpt: Generating Compositional Worlds from Grounded Videos

研究人员提出 WorldSculpt,旨在生成包含数百个物体的杂乱场景的组成式 3D 表示。该方法通过适配强单物体 3D 生成先验至多视图观测,实例化为 Pixal3D,利用多视图条件路径将物体生成锚定在多个姿态观测中。尽管模型仅在标准空间微调单物体,却能泛化至含严重遮挡的大场景而无需场景级训练。此外,团队引入 UE-MeshyScene 基准测试,涵盖数百物体的密集杂乱场景、逐物体标注及真实网格真值。在单物体、受控多物体及 UE-MeshyScene 评估中,该方法表现优于先前的方法,且随场景复杂度和遮挡增加优势更明显。最后,研究展示了将生成的 3DGS 世界(如 Marble 和 HY-World 2.0)转换为组成式网格场景的广泛适用性。

H Hugging Face Papers en 2026-09-07 21:57

Paper page - WorldSculpt: Generating Compositional Worlds from Grounded Videos

The research team proposed WorldSculpt, which achieves scalable combination mesh reconstruction of dense scenes with hundreds of objects by adapting 3D generation priors to multi-view observations for single objects. This method utilizes Pixal3D models and introduces multi-view conditional paths, anchoring object generation in multiple pose observations; although the model is only fine-tuned on single objects in standard space, it can generalize to heavily occluded large scenes without scene-level training. The research team also released the UE-MeshyScene benchmark dataset, which includes real-photo scenes with hundreds of objects, along with per-object annotations and true meshes. Experiments show that this method outperforms existing methods in single-object, multi-object control, and UE-MeshyScene evaluations, and its advantages become more apparent as scene complexity and occlusion increase. Additionally, the study demonstrates the broad applicability of converting generated 3DGS worlds (such as Marble and HY-World 2.0) into combined mesh scenes.