AuraTracer智迹闻
中文

EVENT DOSSIER

From Intent to Evidence: Policy-Steered Multi-Strategy Retrieval for Long-Video Agents

2026-09-07 12:00 Models 🔥 42.2 heat score
1sources
1days unfolding
42.2heat score
5mentions
SummaryAI generated

The research team proposed VESTA, a trained-free long-video intelligent agent system. This system infers focus, recall, or contrast retrieval strategies and evidence account configurations during the exploration phase through an intention router, and employs a routing-based - verification - consolidation cycle for task processing. VESTA uses multimodal evidence manipulation to transform retrieval results into observations, and integrates location, source, conflicts, and verification results through a time evidence ledger to form an adaptive view to guide subsequent actions. In the Video-MME-v2 test, its average accuracy improved by 2.7 points compared to VideoARM; in the LongVideoBench long-subset and LVBench, it improved by 6.9 points and 1.5 points respectively, and performed equally with VideoARM in EgoSchema.

Related eventsRELATED EVENTS
Key entitiesKEY ENTITIES
EgoSchemaLongVideoBenchVESTAVideo-MME-v2VideoARM

Coverage · reports per dayLANGUAGE SPLIT

Entity relations
EgoSchema × LongVideoBe…1EgoSchema × VESTA1EgoSchema × Video-MME-v21EgoSchema × VideoARM1LongVideoBench × VESTA1LongVideoBench × Video-…1

SignalsSIGNALS

Keyword heat
  • VESTA1
  • Video-MME-v21
  • VideoARM1
  • LongVideoBench1
  • EgoSchema1

All reports (1)SOURCES

A arXiv cs.CV en 2026-09-07 12:00

From Intent to Evidence: Policy-Steered Multi-Strategy Retrieval for Long-Video Agents

提出 VESTA,一种无需训练的长视频智能体,通过意图路由器在探索前推断聚焦、召回或对比检索策略及证据账户配置。该智能体采用路由条件获取 - 验证 - 巩固循环,利用多模态证据操作将检索结果转化为观察值,并由时间证据账本整合为包含位置、来源、冲突及验证结果的自适应视图以指导后续行动。在 Video-MME-v2 上,VESTA 平均准确率较 VideoARM 提升 2.7 分;在 LongVideoBench 长子集和 LVBench 上分别提升 6.9 分和 1.5 分,并在 EgoSchema 上与 VideoARM 持平。