Where to Look Matters: Learning Influential Views for VLM-based 3D Visual Grounding
2026-09-07 12:00Models🔥 42.2 heat score
1sources
1days unfolding
42.2heat score
3mentions
SummaryAI generated
To address the issue that existing zero-shot 3D visual localization methods rely on heuristic rules rather than location-related information, researchers proposed the IVGround framework. This framework trains a lightweight view selector to identify views that provide discriminative evidence and uses a two-stage rejection sampling process combined with reasoning VLM feedback to generate supervisory signals. In the reasoning stage, the selector predicts the conditional influential views of each candidate object, which are then evaluated by a frozen reasoning VLM through comparison with the actual location. Experiments show that the IVGround framework consistently improves the localization accuracy of existing zero-shot processes on the ScanRefer and NR3D datasets, demonstrating that “where to look” is crucial for effective 3D visual localization.