ElderBench: Benchmarking Autonomous Mobile Agents for Older Adults
2026-09-07 12:00Science🔥 42.2 heat score
1sources
1days unfolding
42.2heat score
1mentions
SummaryAI generated
On September 7, 2026, the ElderBench study was published on arXiv cs.AI. This is the first benchmark for evaluating mobile graphical interface agents in real-world scenarios for older adults. The study was based on 249 naturally collected smartphone tasks from 20 applications and analyzed the differences between elderly-directed instructions at the syntactic, semantic, and pragmatic levels compared to existing GUI benchmarks. Test results showed that mainstream GUI agents and visual-language models performed significantly worse when handling elderly-directed instructions. By controlling instruction normalization, failure analysis, and detailed language feature analysis, the study identified specific language patterns unique to older adults that caused agents to fail, aiming to provide insights for designing more adaptive, interpretable, and age-friendly GUI agents.