AuraTracer智迹闻
中文

EVENT DOSSIER

MMTClinic: Multimodal, Multilingual Time Series Question Answering and Reasoning Benchmark for Clinical Domain

2026-09-07 12:00 Science 🔥 42.2 heat score
1sources
1days unfolding
42.2heat score
1mentions
SummaryAI generated

The researchers released the MMTClinic benchmark, aimed at evaluating the performance of large language models in clinical time-series question-answering and reasoning tasks. The benchmark integrates multimodal data such as text, medical images, and multi-variable physiological signals, comprising 30,000 pairs of question-answer data in five languages: English, Hindi, Bengali, Marathi, and Tamil (including 15,000 multiple-choice questions and 15,000 open-ended questions). The benchmark covers three core clinical tasks: mortality prediction, heart rate prediction, and SOFA score estimation. The research team evaluated 13 advanced LLMs under zero-sample, few-sample, and thought-chain settings, finding significant differences in model performance across tasks, languages, and modalities, revealing the limitations of current clinical reasoning capabilities. The dataset will be made public after it is accepted for use.

Related eventsRELATED EVENTS
Key entitiesKEY ENTITIES
MMTClinic

SignalsSIGNALS

Keyword heat
  • MMTClinic1

All reports (1)SOURCES

A arXiv cs.AI en 2026-09-07 12:00

MMTClinic: Multimodal, Multilingual Time Series Question Answering and Reasoning Benchmark for Clinical Domain

研究人员发布 MMTClinic 基准,用于评估大语言模型在临床时间序列问答与推理任务上的表现。该基准整合文本、医学图像及多变量生理信号,包含涵盖英语、印地语、孟加拉语、马拉地语和泰米尔语共五种语言的 30,000 对问答数据(含 15,000 道选择题和 15,000 道开放题),覆盖死亡率预测、心率预测及 SOFA 评分估算三项临床任务。研究团队在零样本、少样本及思维链设置下评估了 13 种最先进的 LLM,发现模型在不同任务、语言和模态下的表现存在显著差异,揭示了当前临床推理能力的局限。数据集将在工作被接受后公开。