AuraTracer智迹闻
中文

EVENT DOSSIER

Where to Look Matters: Learning Influential Views for VLM-based 3D Visual Grounding

2026-09-07 12:00 Models 🔥 42.2 heat score
1sources
1days unfolding
42.2heat score
3mentions
SummaryAI generated

To address the issue that existing zero-shot 3D visual localization methods rely on heuristic rules rather than location-related information, researchers proposed the IVGround framework. This framework trains a lightweight view selector to identify views that provide discriminative evidence and uses a two-stage rejection sampling process combined with reasoning VLM feedback to generate supervisory signals. In the reasoning stage, the selector predicts the conditional influential views of each candidate object, which are then evaluated by a frozen reasoning VLM through comparison with the actual location. Experiments show that the IVGround framework consistently improves the localization accuracy of existing zero-shot processes on the ScanRefer and NR3D datasets, demonstrating that “where to look” is crucial for effective 3D visual localization.

Related eventsRELATED EVENTS
Key entitiesKEY ENTITIES
IVSGroundNR3DScanRefer

Coverage · reports per dayLANGUAGE SPLIT

Entity relations
IVSGround × NR3D1IVSGround × ScanRefer1NR3D × ScanRefer1

SignalsSIGNALS

Keyword heat
  • IVSGround1
  • ScanRefer1
  • NR3D1

All reports (1)SOURCES

A arXiv cs.CV en 2026-09-07 12:00

Where to Look Matters: Learning Influential Views for VLM-based 3D Visual Grounding

IVSGround 框架提出了一种用于 VLM 基于的 3D 视觉定位的影响性视图选择方法,旨在解决现有零-shot 方法依赖启发式规则而非定位相关性的问题。该框架通过训练轻量级视图选择器来识别提供判别性证据的视图,并利用两阶段拒绝采样过程结合推理 VLM 反馈生成监督信号。在推理阶段,选择器预测每个候选对象的条件化影响视图,并由冻结的推理 VLM 通过对比定位进行评估。在 ScanRefer 和 NR3D 数据集上的实验表明,IVSGround 一致地提升了现有零-shot 流程的定位精度,证明了“看哪里”对有效 3D 视觉定位至关重要。