RoGe: Novel View Synthesis via End-to-End Implicit Reconstruction and Generation
2026-09-07 12:00Models🔥 42.2 heat score
1sources
1days unfolding
42.2heat score
2mentions
SummaryAI generated
RoGe proposes a unified end-to-end reconstruction and generation framework aimed at synthesizing time-consistent roaming videos using sparse input views. This method constructs an implicit scene representation from a small number of located images through a feedforward reconstruction model, and uses target camera ray queries to obtain geometric features for each frame. These features are then used as conditions in the video diffusion model for generation, without the need for a 3D intermediate representation throughout the process. Experiments on the DL3DV dataset show that RoGe outperforms reconstruction-based, generation-based, and hybrid baseline methods in both image-level metrics and video-level time consistency. A ablation study confirms that implicit features derived from ray queries perform better than original reconstruction tokens or rendered RGB as conditions, and joint training brings further improvements.