AuraTracer智迹闻
中文

EVENT DOSSIER

ICM-Bench: Person-Level Identity Reasoning in Multimodal Agents with Long-Term Memory

2026-09-07 12:00 Models 🔥 42.2 heat score
1sources
1days unfolding
42.2heat score
2mentions
SummaryAI generated

ICM-Bench, as the first benchmark for evaluating multi-modal agents in long-video identity centralization reasoning, includes 839 synthetic segments, a duration of 141 minutes, and 1,217 open-ended questions for six adults. Compared with the direct caption-memory baseline, memory-enhancing agents, and graph retrieval systems, Gemini 3.1 Pro achieved an overall accuracy of 74.0%, but dropped to 60.3% on questions requiring long-term identity profiles. The results show that while current multi-modal systems can restore much event-level memory, their reliability is low when evidence needs to be accumulated around stable characters.

Related eventsRELATED EVENTS
Key entitiesKEY ENTITIES
Gemini 3.1 ProICM-Bench

Coverage · reports per dayLANGUAGE SPLIT

Entity relations
Gemini 3.1 Pro × ICM-Be…1

SignalsSIGNALS

Keyword heat
  • ICM-Bench1
  • Gemini 3.1 Pro1

All reports (1)SOURCES

A arXiv cs.CV en 2026-09-07 12:00

ICM-Bench: Person-Level Identity Reasoning in Multimodal Agents with Long-Term Memory

ICM-Bench is the first benchmark specifically designed to evaluate multi-modal agents in long-video memory for identity-centered reasoning. The benchmark includes 839 synthetic segments, a duration of 141 minutes, and 1,217 open-ended questions about six adults. Compared with the direct caption-memory baseline, memory-enhancing agents, and graph retrieval systems, Gemini 3.1 Pro achieved an overall accuracy of 74.0%, but dropped to 60.3% on questions requiring long-term identity profiles. The results indicate that while current systems can recover much event-level memory, their reliability is low when evidence needs to be accumulated around stable characters.