ICM-Bench: Person-Level Identity Reasoning in Multimodal Agents with Long-Term Memory
2026-09-07 12:00Models🔥 42.2 heat score
1sources
1days unfolding
42.2heat score
2mentions
SummaryAI generated
ICM-Bench, as the first benchmark for evaluating multi-modal agents in long-video identity centralization reasoning, includes 839 synthetic segments, a duration of 141 minutes, and 1,217 open-ended questions for six adults. Compared with the direct caption-memory baseline, memory-enhancing agents, and graph retrieval systems, Gemini 3.1 Pro achieved an overall accuracy of 74.0%, but dropped to 60.3% on questions requiring long-term identity profiles. The results show that while current multi-modal systems can restore much event-level memory, their reliability is low when evidence needs to be accumulated around stable characters.
ICM-Bench is the first benchmark specifically designed to evaluate multi-modal agents in long-video memory for identity-centered reasoning. The benchmark includes 839 synthetic segments, a duration of 141 minutes, and 1,217 open-ended questions about six adults. Compared with the direct caption-memory baseline, memory-enhancing agents, and graph retrieval systems, Gemini 3.1 Pro achieved an overall accuracy of 74.0%, but dropped to 60.3% on questions requiring long-term identity profiles. The results indicate that while current systems can recover much event-level memory, their reliability is low when evidence needs to be accumulated around stable characters.