From Intent to Evidence: Policy-Steered Multi-Strategy Retrieval for Long-Video Agents
2026-09-07 12:00Models🔥 42.2 heat score
1sources
1days unfolding
42.2heat score
5mentions
SummaryAI generated
The research team proposed VESTA, a trained-free long-video intelligent agent system. This system infers focus, recall, or contrast retrieval strategies and evidence account configurations during the exploration phase through an intention router, and employs a routing-based - verification - consolidation cycle for task processing. VESTA uses multimodal evidence manipulation to transform retrieval results into observations, and integrates location, source, conflicts, and verification results through a time evidence ledger to form an adaptive view to guide subsequent actions. In the Video-MME-v2 test, its average accuracy improved by 2.7 points compared to VideoARM; in the LongVideoBench long-subset and LVBench, it improved by 6.9 points and 1.5 points respectively, and performed equally with VideoARM in EgoSchema.