Development and validation of an interpretable machine learning model for predicting in-hospital mortality in patients with pulmonary fibrosis
In-hospital mortality risk varies substantially among patients with pulmonary fibrosis (PF). Existing prognostic models have primarily focused on long-term survival and have limited ability to capture early in-hospital mortality risk or incorporate worsening oxygenation, inflammatory burden, and treatment-escalation information. This study aimed to develop and internally validate an interpretable machine learning model for in-hospital mortality risk stratification using early admission clinical data and information on ventilatory support within 24 h after admission in hospitalized patients with PF. This retrospective cohort study included hospitalized patients with PF admitted to a tertiary hospital between January 2019 and December 2023. Candidate variables included demographic characteristics, comorbidities, disease subtype, clinical symptoms, laboratory tests, arterial blood gas analysis, modified Medical Research Council (mMRC) dyspnea score, and the use of invasive or noninvasive ventilation within 24 h after admission. Nested LASSO feature selection combined with repeated 10-fold cross-validation was used for model development and internal validation. The main candidate models included logistic regression (LR), naive Bayes (NB), support vector machine (SVM), random forest (RF), neural network (NN), multilayer perceptron (MLP), adaptive boosting (AdaBoost), and gradient boosting machine (GBM). A supplementary clinically simplified model based on LASSO prescreening, recursive feature elimination, and logistic regression (RFE_LR) was also developed. Model performance was evaluated using the area under the receiver operating characteristic curve (AUC), threshold-dependent classification metrics, calibration analysis, decision curve analysis, clinical impact curves, bootstrap internal validation, and SHAP interpretation. Additional analyses included pairwise AUC comparisons with Holm adjustment for multiple testing, alternative-threshold analyses, exploratory in-sample recalibration analyses, and sensitivity analyses related to noninvasive ventilation (NIV). A total of 547 patients were included, of whom 108 met the prespecified primary mortality outcome, corresponding to an event rate of 19.74%. Six variables met the prespecified ≥ 50% selection-frequency criterion across the outer cross-validation folds: neutrophil percentage (NE%), NIV use within 24 h after admission, hypertension, mMRC score, idiopathic pulmonary fibrosis (IPF), and arterial oxygen partial pressure (PaO₂). Among the main candidate models, the NN model achieved the highest patient-level aggregated repeated-cross-validation AUC of 0.821 (nominal 95% CI: 0.772–0.869) and was selected as the final main model. This estimate was calculated after averaging each patient’s out-of-fold predicted probabilities across the 10 cross-validation repeats. After Holm adjustment for eight pairwise comparisons, only the AUC difference between the NN and NB models remained statistically significant (Holm-adjusted P = 0.040), whereas all other pairwise differences were not statistically significant. RFE_LR achieved an AUC of 0.789 (95% CI: 0.734–0.844), suggesting that the simplified interpretable model retained complementary value. At the maximum-F1 threshold of 0.57, the final NN model had a sensitivity of 0.602 and a specificity of 0.897. Lower thresholds improved death detection but reduced specificity. Calibration analysis showed that the model reflected mortality risk gradients, although raw predicted probabilities were systematically overestimated. The NIV-only AUC was 0.626. In the clinically relevant NIV = 0 subgroup, the final NN model achieved an AUC of 0.769; however, at the primary threshold of 0.57, sensitivity was only 0.456 and specificity was 0.912. After NIV was excluded and the models were redeveloped, the highest AUC was 0.790. These findings indicate that the model retained discriminative information beyond NIV status; however, the low sensitivity at the primary threshold limits its use as a stand-alone early-warning rule in patients not receiving NIV. Bootstrap internal validation yielded an optimism-corrected AUC of 0.788 and a mean OOB AUC of 0.718, with a 2.5th–97.5th percentile range of 0.604–0.817. Across the principal internal-validation approaches, AUC point estimates ranged from 0.718 to 0.821, indicating moderate but method-dependent discrimination. SHAP analysis identified mMRC score, IPF, PaO₂, NE%, hypertension, and NIV use within 24 h after admission as the main contributing variables. This study developed and internally validated an interpretable machine learning model for early in-hospital mortality risk stratification in hospitalized patients with PF. The model showed moderate, method-dependent discrimination and may be more appropriate for risk ranking than for stand-alone death detection or direct absolute-risk estimation. Performance differences among most evaluated algorithms were limited, and raw predicted probabilities were systematically overestimated. In the NIV = 0 subgroup, the primary threshold showed high specificity but low sensitivity, further limiting its use as a stand-alone early-warning rule. External validation and appropriate recalibration are required before clinical implementation.
Authors
- Yadong Yuan (ORCID: https://orcid.org/0000-0002-1319-4743)
- Pei Wang (ORCID: https://orcid.org/0000-0002-3586-5130)
- Xiaodan Jiao
- Yanping Zhang
Institutions
- Hebei Medical University (CN)
- Second Hospital of Hebei Medical University (CN)
Publication Details
- Journal
- BMC Medical Informatics and Decision Making
- Published
- 2026-09-04
- DOI
- https://doi.org/10.1186/s12911-026-03814-5
- Primary Topic
- Interstitial Lung Diseases and Idiopathic Pulmonary Fibrosis
- Type
- article
- Field-Weighted Citation Impact
- 0.00