Cross-modal triage network: a multimodal deep learning framework for severity-based triage and visual explainability in chest radiographs
2026-09-07 12:00Science🔥 42.2 heat score
1sources
1days unfolding
42.2heat score
5mentions
SummaryAI generated
The researchers developed the Cross-Modal Diagnosis Network (CMTN), a multimodal deep learning framework that combines images and text, aimed at diagnosing chest X-rays based on severity, performing pathological analysis, and providing visual interpretation. This network integrates the Swin Transformer V2 visual encoder and the PubMedBERT text encoder, and was trained on the MIMIC-CXR-JPG dataset. It was optimized for level 4 severity diagnosis and 14 common pathologies. Quantitative benchmark tests showed that CMTN had a weighted Kappa coefficient of 0.9341 in severity diagnosis, a macro AUROC of 0.9970, and processing latency of only 34 milliseconds, which is significantly better than the BioViL baseline model. However, blind clinical audits revealed that the model’s consistency with the judgments of real radiologists was significantly reduced (weighted Kappa coefficient: 0.1399), and only 54.3% of the heatmaps met acceptable quality standards.