AuraTracer智迹闻
中文

EVENT DOSSIER

XDG: Accelerated Visual Disambiguation

2026-09-07 12:00 Models 🔥 42.2 heat score
1sources
1days unfolding
42.2heat score
2mentions
SummaryAI generated

The XDG model fine-tunes Depth Anything 3 using a lightweight LoRA adapter, reusing the camera token as a paired classification token and utilizing its cross-view geometric reasoning capabilities for visual disambiguation. This approach avoids the additional computational overhead of heavy decoders. Experiments show that XDG remains competitive with state-of-the-art methods in paired and reconstruction benchmarks, while providing more than three times faster inference speeds; in a single LaMAR scenario involving thousands of images, it can save over 10 hours of time spent on visual disambiguation. The related code is now open-source.

Related eventsRELATED EVENTS
Key entitiesKEY ENTITIES
Depth Anything 3XDG

Event frameEVENT FRAME

Launch

arXiv:2608.29733v2 XDG 提出高效视觉去重模型,推理速度提升 3 倍

Coverage · reports per dayLANGUAGE SPLIT

Entity relations
Depth Anything 3 × XDG1

SignalsSIGNALS

Keyword heat
  • XDG1
  • Depth Anything 31

All reports (1)SOURCES

A arXiv cs.CV en 2026-09-07 12:00

XDG: Accelerated Visual Disambiguation

XDG 模型通过轻量级 LoRA 适配器微调 Depth Anything 3,将相机令牌重用作成对分类令牌,实现了视觉去歧义的高效处理。该方法利用 3D 基础模型固有的跨视图几何推理能力,直接适配骨干网络表示,避免了重型解码器的额外计算开销。实验表明,XDG 在成对和重建基准测试中与最先进方法保持竞争力,同时提供超过 3 倍的推理速度提升;在包含数千张图像的单个 LaMAR 场景中,可节省超过 10 小时的视觉去歧义处理时间。代码已开源。