AuraTracer智迹闻
中文

EVENT DOSSIER

ElderBench: Benchmarking Autonomous Mobile Agents for Older Adults

2026-09-07 12:00 Science 🔥 42.2 heat score
1sources
1days unfolding
42.2heat score
1mentions
SummaryAI generated

On September 7, 2026, the ElderBench study was published on arXiv cs.AI. This is the first benchmark for evaluating mobile graphical interface agents in real-world scenarios for older adults. The study was based on 249 naturally collected smartphone tasks from 20 applications and analyzed the differences between elderly-directed instructions at the syntactic, semantic, and pragmatic levels compared to existing GUI benchmarks. Test results showed that mainstream GUI agents and visual-language models performed significantly worse when handling elderly-directed instructions. By controlling instruction normalization, failure analysis, and detailed language feature analysis, the study identified specific language patterns unique to older adults that caused agents to fail, aiming to provide insights for designing more adaptive, interpretable, and age-friendly GUI agents.

Related eventsRELATED EVENTS
Key entitiesKEY ENTITIES
ElderBench

SignalsSIGNALS

Keyword heat
  • ElderBench1

All reports (1)SOURCES

A arXiv cs.AI en 2026-09-07 12:00

ElderBench: Benchmarking Autonomous Mobile Agents for Older Adults

ElderBench 是首个针对老年人真实场景评估移动图形界面代理的基准,由来自 20 个应用的 249 项自然采集智能手机任务构建。该研究分析了老年指令与现有 GUI 基准在句法、语义及语用层面的语言差异,并在线下及线上环境中测试主流 GUI 代理和视觉 - 语言模型,发现处理老年导向指令时性能显著下降。通过控制指令归一化、失败分析及细粒度语言特征分析,识别了特定于老年人的语言模式导致代理失败的原因,旨在为更自适应、可解释且包容年龄的 GUI 代理提供设计洞察。