An interpretable Day-1 machine-learning model for COVID-19 mortality and severity prognosis: development in a Pakistani hospital cohort and external evaluation in Chinese and Italian populations

Early risk stratification of COVID-19 patients at admission informs triage and resource allocation, particularly in low- and middle-income settings where intensive-care capacity is constrained. Most published models are not interpretable, rarely externally validated, and developed almost exclusively in high-income cohorts. We developed an interpretable random-forest classifier for three-class COVID-19 severity (Mild: not ventilated, survived; Severe: ventilated, survived; Fatal: died) using only data available within 24 hours of admission in 321 hospitalised PCR-confirmed COVID-19 patients in Pakistan. From 2,278 raw variables we retained 20 predictors by clinical screening and permutation-importance selection, with leakage-controlled multiple imputation by chained equations and SHapley Additive exPlanations (SHAP) for interpretation. A 15-feature variant was evaluated in two independent public cohorts: the Chinese iCTCF cohort (n = 894) and an Italian complete-blood-count cohort (n = 1,218). The development model achieved a macro F1 of 0.55, identifying the Fatal class most reliably (F1 0.67, recall 0.73) and Mild cases least well. SHAP identified a parsimonious set of admission-time predictors (oxygen saturation, lactate dehydrogenase, urea, C-reactive protein, age and radiographic pneumonia), including a raised urea-to-creatinine ratio in fatal cases most consistent with dehydration (median urea 80 vs 50 mg/dL, fatal vs mild; p < 0.001). Externally, the model retained useful discrimination in the Chinese cohort (mortality AUC 0.82 in 719 patients with recorded outcomes; severe-or-worse AUC 0.75; macro F1 0.50) but only modest discrimination in the Italian cohort (AUC 0.62, 95% CI 0.56-0.68), which we treat as a feature-availability stress test rather than a second validation. Cross-cohort SHAP rankings correlated strongly in China (Spearman ρ = 0.91) and moderately in Italy (ρ = 0.76), where most laboratory predictors were imputed and carried no information beyond the complete blood count; discrimination tracked how many predictors were actually measured. The admission-time signature converges with established markers of hypoperfusion and inflammation.

Authors

Institutions

Publication Details

Journal
medRxiv
Published
2026-10-05
DOI
https://doi.org/10.64898/2026.10.01.26364553
Primary Topic
COVID-19 Clinical Research Studies
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
preprint

An interpretable Day-1 machine-learning model for COVID-19 mortality and severity prognosis: development in a Pakistani hospital cohort and external evaluation in Chinese and Italian populations

Shaper Mirza, Tehmina Mustafa, Nida Javaid, Muhammad Adeel Zaffar et al.
medRxiv
COVID-19 Clinical Research Studies
preprint

An interpretable Day-1 machine-learning model for COVID-19 mortality and severity prognosis: development in a Pakistani hospital cohort and external evaluation in Chinese and Italian populations

Shaper Mirza, Tehmina Mustafa, Nida Javaid, Muhammad Adeel Zaffar, Aasia Khaliq, Eesha Tariq, Talha Jaffer, Shakir Aslam
preprint en

Abstract

Early risk stratification of COVID-19 patients at admission informs triage and resource allocation, particularly in low- and middle-income settings where intensive-care capacity is constrained. Most published models are not interpretable, rarely externally validated, and developed almost exclusively in high-income cohorts. We developed an interpretable random-forest classifier for three-class COVID-19 severity (Mild: not ventilated, survived; Severe: ventilated, survived; Fatal: died) using only data available within 24 hours of admission in 321 hospitalised PCR-confirmed COVID-19 patients in Pakistan. From 2,278 raw variables we retained 20 predictors by clinical screening and permutation-importance selection, with leakage-controlled multiple imputation by chained equations and SHapley Additive exPlanations (SHAP) for interpretation. A 15-feature variant was evaluated in two independent public cohorts: the Chinese iCTCF cohort (n = 894) and an Italian complete-blood-count cohort (n = 1,218). The development model achieved a macro F1 of 0.55, identifying the Fatal class most reliably (F1 0.67, recall 0.73) and Mild cases least well. SHAP identified a parsimonious set of admission-time predictors (oxygen saturation, lactate dehydrogenase, urea, C-reactive protein, age and radiographic pneumonia), including a raised urea-to-creatinine ratio in fatal cases most consistent with dehydration (median urea 80 vs 50 mg/dL, fatal vs mild; p < 0.001). Externally, the model retained useful discrimination in the Chinese cohort (mortality AUC 0.82 in 719 patients with recorded outcomes; severe-or-worse AUC 0.75; macro F1 0.50) but only modest discrimination in the Italian cohort (AUC 0.62, 95% CI 0.56-0.68), which we treat as a feature-availability stress test rather than a second validation. Cross-cohort SHAP rankings correlated strongly in China (Spearman ρ = 0.91) and moderately in Italy (ρ = 0.76), where most laboratory predictors were imputed and carried no information beyond the complete blood count; discrimination tracked how many predictors were actually measured. The admission-time signature converges with established markers of hypoperfusion and inflammation.

medRxiv
Lahore University of Management Sciences (PK), Wellcome Sanger Institute (GB), University of Bergen (NO)
Good health and well-being
COVID-19 Clinical Research Studies
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.