Cross-dataset transportability of pediatric chest X-ray deep learning across three countries: discrimination, calibration, operating-point failure, and limited-label recovery
A cross-dataset deep learning study on pediatric pneumonia classification evaluated the performance of imaging data from three countries: Guangzhou, Bangladesh, and Vietnam. The research team conducted internal testing using 5,824 radiological images from Guangzhou, and the model achieved an AUROC of 0.976 on the source data, with a sensitivity of 95.1%. However, when zero-samples were migrated to the Bangladesh BDCXR dataset (3,257 cases) and the Vietnam VinDr-PCXR/PediCXR dataset (1,077 cases), performance significantly declined: the AUROC dropped to 0.798 and 0.742 respectively, and sensitivity under a fixed threshold fell to 6.2% and 0%, indicating severe operational point failures. Limited label restoration experiments on the Bangladesh data showed that using 163 labels for Platt recalibration could increase sensitivity to 88.3%, but specificity was only 47.9%. Repeated fitting confirmed that sensitivity recovery was accompanied by significant fluctuations in specificity. The study revealed cross-national data differences in discrimination ability, calibration, and operational points…