MMTClinic: Multimodal, Multilingual Time Series Question Answering and Reasoning Benchmark for Clinical Domain
2026-09-07 12:00Science🔥 42.2 heat score
1sources
1days unfolding
42.2heat score
1mentions
SummaryAI generated
The researchers released the MMTClinic benchmark, aimed at evaluating the performance of large language models in clinical time-series question-answering and reasoning tasks. The benchmark integrates multimodal data such as text, medical images, and multi-variable physiological signals, comprising 30,000 pairs of question-answer data in five languages: English, Hindi, Bengali, Marathi, and Tamil (including 15,000 multiple-choice questions and 15,000 open-ended questions). The benchmark covers three core clinical tasks: mortality prediction, heart rate prediction, and SOFA score estimation. The research team evaluated 13 advanced LLMs under zero-sample, few-sample, and thought-chain settings, finding significant differences in model performance across tasks, languages, and modalities, revealing the limitations of current clinical reasoning capabilities. The dataset will be made public after it is accepted for use.