AuraTracer智迹闻
中文

EVENT DOSSIER

Cross-Domain Tracker Adaptation Without Target-Domain Labels via Vision-Language Agents

2026-09-07 12:00 Models across 2 days 🔥 47.2 heat score
2sources
2days unfolding
47.2heat score
2mentions
SummaryAI generated

The researchers proposed a cross-domain tracker adaptation system that does not require target domain labels, using a visual language model (VLM) as a diagnostic agent. This method performed well in the experiments ranging from MOT17 to MOT20, improving the average HOTA from 0.267 in the source domain configuration to nearly the upper limit of the target domain, at 0.357. Specifically, the average recovery loss performance increased by 67.8%, with a recovery rate of 86.7% in the highest-density sequences. The system directly checks the rendering output through an iterative tuning loop and identifies visual failure patterns, adjusting parameters only when clear failures are detected, thereby overcoming the vulnerability of traditional label-free optimization methods to large domain offsets. The study indicates that the effectiveness of this method depends on whether domain offsets are reflected through detection-level parameter exposure; its effectiveness is limited in scenarios where the source domain configuration is already close to optimal (such as MOT17 to DanceTrack).

Related eventsRELATED EVENTS
Key entitiesKEY ENTITIES
VLMVision-Language Model

Coverage · reports per dayLANGUAGE SPLIT

Entity relations
VLM × Vision-Language M…1

Integrated timelineUNIFIED TIMELINE

  1. 2026-09-04

    Cross-Domain Tracker Adaptation Without…

    研究人员提出一种无需目标域标签的跨域跟踪器适配系统,利用视觉语言模型(VLM)作为诊断代理。在 MOT17 至 MOT20 的实验中,该方法将平均 HOTA 从源域配置下的 0.267 提升至接近目标域上限的 0.357,恢复损失性能达 …

  2. 2026-09-07

    Cross-Domain Tracker Adaptation Without…

    研究人员提出一种无需目标域标签的跨域跟踪器适配系统,利用视觉语言模型(VLM)作为诊断代理。在 MOT17 至 MOT20 的实验中,该方法将平均 HOTA 从源域配置下的 0.267 提升至接近目标域上限的 0.357,恢复损失性能达 …

SignalsSIGNALS

Keyword heat
  • Vision-Language Model1
  • VLM1

All reports (2)SOURCES

A arXiv cs.CV en 2026-09-04 23:05

Cross-Domain Tracker Adaptation Without Target-Domain Labels via Vision-Language Agents

研究人员提出一种无需目标域标签的跨域跟踪器适配系统,利用视觉语言模型(VLM)作为诊断代理。在 MOT17 至 MOT20 的实验中,该方法将平均 HOTA 从源域配置下的 0.267 提升至接近目标域上限的 0.357,恢复损失性能达 67.8%,最高密度序列下恢复率高达 86.7%。该系统通过迭代调优循环直接检查渲染输出并推荐参数更新,克服了传统基于标注指标的优化在跨域迁移中的脆弱性。与易受大领域偏移影响的手动代理目标贝叶斯优化不同,该 VLM 调优器具有选择性:仅在识别出明确视觉故障模式时修改配置,否则保持原状以保留简单迁移的性能优势。该方法的有效性取决于领域偏移是否通过暴露的检测级参数体现,在 MOT17 至 DanceTrack 等源域已接近最优的场景中效果有限。

A arXiv cs.CV en 2026-09-07 12:00

Cross-Domain Tracker Adaptation Without Target-Domain Labels via Vision-Language Agents

研究人员提出一种无需目标域标签的跨域跟踪器适配系统,利用视觉语言模型(VLM)作为诊断代理。在 MOT17 至 MOT20 的实验中,该方法将平均 HOTA 从源域配置下的 0.267 提升至接近目标域上限的 0.357,恢复损失性能达 67.8%,最高密度序列下恢复率达 86.7%。与易受大领域偏移影响的无标签贝叶斯优化不同,该 VLM 调优器通过迭代循环直接检查渲染输出并识别视觉失效模式,仅在检测到明确故障时调整参数,从而在简单迁移中保持性能的同时改善困难案例。研究还指出该方法在检测级参数暴露的领域偏移下有效,而在源域配置已接近最优(如 MOT17 至 DanceTrack)的场景中效果有限。