In September 2026, researchers launched WearableQA, a benchmark dataset designed to evaluate the health reasoning capabilities of large language models. The dataset contains 4,084 multiple-choice questions, built based on 500 days of wearable device data, blood biomarkers, and demographic information from 200 real users. WearableQA uses a dual grounding framework, combining literature-based and statistical validation patterns, while preserving real-world distribution characteristics such as device noise and individual differences. It includes 16 question types covering data reasoning and health reasoning, as well as single-signal and cross-signal reasoning. Assessments of 14 proprietary and open-source large language models showed that their performance ranged from 19.6% to 72.9%, far exceeding the 10% random baseline. Most models had an accuracy below 60%, indicating that this benchmark can effectively distinguish model capabilities and demonstrate that the field of health reasoning still has unresolved issues.