AuraTracer智迹闻
中文

EVENT DOSSIER

WearableQA: A Benchmark for Health Reasoning over Real-World Wearable Data

2026-09-07 12:00 Science across 2 days 🔥 47.2 heat score
2sources
2days unfolding
47.2heat score
1mentions
SummaryAI generated

In September 2026, researchers launched WearableQA, a benchmark dataset designed to evaluate the health reasoning capabilities of large language models. The dataset contains 4,084 multiple-choice questions, built based on 500 days of wearable device data, blood biomarkers, and demographic information from 200 real users. WearableQA uses a dual grounding framework, combining literature-based and statistical validation patterns, while preserving real-world distribution characteristics such as device noise and individual differences. It includes 16 question types covering data reasoning and health reasoning, as well as single-signal and cross-signal reasoning. Assessments of 14 proprietary and open-source large language models showed that their performance ranged from 19.6% to 72.9%, far exceeding the 10% random baseline. Most models had an accuracy below 60%, indicating that this benchmark can effectively distinguish model capabilities and demonstrate that the field of health reasoning still has unresolved issues.

Related eventsRELATED EVENTS
Key entitiesKEY ENTITIES
WearableQA

Integrated timelineUNIFIED TIMELINE

  1. 2026-09-04

    WearableQA: A Benchmark for Health Reas…

    研究人员推出 WearableQA,该基准包含由 200 名真实用户数据(含最长 500 天测量记录)构建的 4,084 道选择题。该基准涵盖可穿戴时间序列、血液生物标志物及人口统计学信息,并保留设备噪声与个体差异等真实分布特征。为评估不…

  2. 2026-09-07

    WearableQA: A Benchmark for Health Reas…

    研究人员推出 WearableQA,这是一个包含 4,084 道选择题的基准测试数据集,基于 200 名真实用户长达 500 天的可穿戴设备时间序列、血液生物标志物及人口统计数据构建。该基准通过双 grounding 框架结合文献与统计验…

SignalsSIGNALS

Keyword heat
  • WearableQA2

All reports (2)SOURCES

A arXiv cs.CL en 2026-09-05 01:52

WearableQA: A Benchmark for Health Reasoning over Real-World Wearable Data

研究人员推出 WearableQA,该基准包含由 200 名真实用户数据(含最长 500 天测量记录)构建的 4,084 道选择题。该基准涵盖可穿戴时间序列、血液生物标志物及人口统计学信息,并保留设备噪声与个体差异等真实分布特征。为评估不同推理能力,研究设计了 16 种题型,按数据与健康推理、单信号与跨信号推理两个维度分类。构建过程采用结合文献生理发现与统计验证人群模式的“双重 grounding"框架。对 14 个专有及开源大语言模型的测试显示,其表现介于 19.6% 至 72.9% 之间(基准为 10%),且多数模型准确率低于 60%,表明该基准尚未被完全解决。WearableQA 旨在提供评估大语言模型在真实世界可穿戴数据上进行健康推理的可靠诊断基准。

A arXiv cs.CL en 2026-09-07 12:00

WearableQA: A Benchmark for Health Reasoning over Real-World Wearable Data

研究人员推出 WearableQA,这是一个包含 4,084 道选择题的基准测试数据集,基于 200 名真实用户长达 500 天的可穿戴设备时间序列、血液生物标志物及人口统计数据构建。该基准通过双 grounding 框架结合文献与统计验证模式,保留了设备噪声和个体差异等真实分布特征,并设计了涵盖数据推理与健康推理、单信号与跨信号推理的 16 种问题类型。对 14 个专有及开源大语言模型的评估显示,其表现范围从 19.6% 至 72.9%,远超 10% 随机基准,且多数模型准确率低于 60%,表明该基准能有效区分模型能力并证明健康推理领域尚未解决。