AuraTracer智迹闻
中文

EVENT DOSSIER

Intrinsic Temporal Adaptation of CLIP for Partially Relevant Video Retrieval

2026-09-07 12:00 Models 🔥 42.2 heat score
1sources
1days unfolding
42.2heat score
1mentions
SummaryAI generated

To address the problems in Partial Relevant Video Retrieval (PRVR), the research team proposed a new framework called Intrinsic Temporal Adaptation (ITA). This framework uses the Backbone-Internal Temporal Adaptation mechanism, enabling the last layer of the visual Transformer to focus on adjacent frames while keeping the main components of the CLIP model frozen, with only the adaptation parameters trained. Additionally, the framework introduces the Affinity-Weighted Gradient Propagation method, which aggregates the first k frames based on the affinity between text and frames to propagate learning signals. Experiments show that this method achieves the highest performance in PRVR benchmarks, demonstrating robustness across different datasets and enabling more accurate frame-level evidence retrieval at relevant times. The related code is now open-source.

Related eventsRELATED EVENTS
Key entitiesKEY ENTITIES
CLIP

SignalsSIGNALS

Keyword heat
  • CLIP1

All reports (1)SOURCES

A arXiv cs.CV en 2026-09-07 12:00

Intrinsic Temporal Adaptation of CLIP for Partially Relevant Video Retrieval

本文提出 Intrinsic Temporal Adaptation (ITA) 框架,用于解决部分相关视频检索(PRVR)问题。该框架通过 Backbone-Internal Temporal Adaptation 使视觉 Transformer 最后几层关注邻接帧组,同时保持 CLIP 冻结仅训练适应参数;并引入 Affinity-Weighted Gradient Propagation 基于文本 - 帧亲和力聚合前 k 帧以传播学习信号。该方法在 PRVR 基准测试中达到最先进水平,展现跨数据集鲁棒性,并在真实相关时刻内检索到更准确的帧级证据。代码已开源。