AuraTracer智迹闻
中文

EVENT DOSSIER

RoGe: Novel View Synthesis via End-to-End Implicit Reconstruction and Generation

2026-09-07 12:00 Models 🔥 42.2 heat score
1sources
1days unfolding
42.2heat score
2mentions
SummaryAI generated

RoGe proposes a unified end-to-end reconstruction and generation framework aimed at synthesizing time-consistent roaming videos using sparse input views. This method constructs an implicit scene representation from a small number of located images through a feedforward reconstruction model, and uses target camera ray queries to obtain geometric features for each frame. These features are then used as conditions in the video diffusion model for generation, without the need for a 3D intermediate representation throughout the process. Experiments on the DL3DV dataset show that RoGe outperforms reconstruction-based, generation-based, and hybrid baseline methods in both image-level metrics and video-level time consistency. A ablation study confirms that implicit features derived from ray queries perform better than original reconstruction tokens or rendered RGB as conditions, and joint training brings further improvements.

Related eventsRELATED EVENTS
Key entitiesKEY ENTITIES
DL3DVRoGe

Coverage · reports per dayLANGUAGE SPLIT

Entity relations
DL3DV × RoGe1

SignalsSIGNALS

Keyword heat
  • RoGe1
  • DL3DV1

All reports (1)SOURCES

A arXiv cs.CV en 2026-09-07 12:00

RoGe: Novel View Synthesis via End-to-End Implicit Reconstruction and Generation

RoGe 提出了一种端到端统一的重建与生成框架,旨在利用稀疏输入视图合成具有时间一致性的漫游视频。该方法通过前馈重建模型从少量已定位图像构建隐式场景表示,并利用目标相机射线查询以获取每帧几何特征,随后将这些特征作为条件注入视频扩散模型进行生成,全程无需 3D 中间表示。RoGe 在 DL3DV 数据集上的实验表明,其在图像级指标和视频级时间一致性方面均优于基于重建、基于生成及混合基线方法;消融研究证实,射线查询的隐式特征表现优于原始重建 token 或渲染 RGB 作为条件,且联合训练带来了进一步增益。