Development and validation of an interpretable machine learning model for predicting in-hospital mortality in patients with pulmonary fibrosis

In-hospital mortality risk varies substantially among patients with pulmonary fibrosis (PF). Existing prognostic models have primarily focused on long-term survival and have limited ability to capture early in-hospital mortality risk or incorporate worsening oxygenation, inflammatory burden, and treatment-escalation information. This study aimed to develop and internally validate an interpretable machine learning model for in-hospital mortality risk stratification using early admission clinical data and information on ventilatory support within 24 h after admission in hospitalized patients with PF. This retrospective cohort study included hospitalized patients with PF admitted to a tertiary hospital between January 2019 and December 2023. Candidate variables included demographic characteristics, comorbidities, disease subtype, clinical symptoms, laboratory tests, arterial blood gas analysis, modified Medical Research Council (mMRC) dyspnea score, and the use of invasive or noninvasive ventilation within 24 h after admission. Nested LASSO feature selection combined with repeated 10-fold cross-validation was used for model development and internal validation. The main candidate models included logistic regression (LR), naive Bayes (NB), support vector machine (SVM), random forest (RF), neural network (NN), multilayer perceptron (MLP), adaptive boosting (AdaBoost), and gradient boosting machine (GBM). A supplementary clinically simplified model based on LASSO prescreening, recursive feature elimination, and logistic regression (RFE_LR) was also developed. Model performance was evaluated using the area under the receiver operating characteristic curve (AUC), threshold-dependent classification metrics, calibration analysis, decision curve analysis, clinical impact curves, bootstrap internal validation, and SHAP interpretation. Additional analyses included pairwise AUC comparisons with Holm adjustment for multiple testing, alternative-threshold analyses, exploratory in-sample recalibration analyses, and sensitivity analyses related to noninvasive ventilation (NIV). A total of 547 patients were included, of whom 108 met the prespecified primary mortality outcome, corresponding to an event rate of 19.74%. Six variables met the prespecified ≥ 50% selection-frequency criterion across the outer cross-validation folds: neutrophil percentage (NE%), NIV use within 24 h after admission, hypertension, mMRC score, idiopathic pulmonary fibrosis (IPF), and arterial oxygen partial pressure (PaO₂). Among the main candidate models, the NN model achieved the highest patient-level aggregated repeated-cross-validation AUC of 0.821 (nominal 95% CI: 0.772–0.869) and was selected as the final main model. This estimate was calculated after averaging each patient’s out-of-fold predicted probabilities across the 10 cross-validation repeats. After Holm adjustment for eight pairwise comparisons, only the AUC difference between the NN and NB models remained statistically significant (Holm-adjusted P = 0.040), whereas all other pairwise differences were not statistically significant. RFE_LR achieved an AUC of 0.789 (95% CI: 0.734–0.844), suggesting that the simplified interpretable model retained complementary value. At the maximum-F1 threshold of 0.57, the final NN model had a sensitivity of 0.602 and a specificity of 0.897. Lower thresholds improved death detection but reduced specificity. Calibration analysis showed that the model reflected mortality risk gradients, although raw predicted probabilities were systematically overestimated. The NIV-only AUC was 0.626. In the clinically relevant NIV = 0 subgroup, the final NN model achieved an AUC of 0.769; however, at the primary threshold of 0.57, sensitivity was only 0.456 and specificity was 0.912. After NIV was excluded and the models were redeveloped, the highest AUC was 0.790. These findings indicate that the model retained discriminative information beyond NIV status; however, the low sensitivity at the primary threshold limits its use as a stand-alone early-warning rule in patients not receiving NIV. Bootstrap internal validation yielded an optimism-corrected AUC of 0.788 and a mean OOB AUC of 0.718, with a 2.5th–97.5th percentile range of 0.604–0.817. Across the principal internal-validation approaches, AUC point estimates ranged from 0.718 to 0.821, indicating moderate but method-dependent discrimination. SHAP analysis identified mMRC score, IPF, PaO₂, NE%, hypertension, and NIV use within 24 h after admission as the main contributing variables. This study developed and internally validated an interpretable machine learning model for early in-hospital mortality risk stratification in hospitalized patients with PF. The model showed moderate, method-dependent discrimination and may be more appropriate for risk ranking than for stand-alone death detection or direct absolute-risk estimation. Performance differences among most evaluated algorithms were limited, and raw predicted probabilities were systematically overestimated. In the NIV = 0 subgroup, the primary threshold showed high specificity but low sensitivity, further limiting its use as a stand-alone early-warning rule. External validation and appropriate recalibration are required before clinical implementation.

Authors

Institutions

Publication Details

Journal
BMC Medical Informatics and Decision Making
Published
2026-09-04
DOI
https://doi.org/10.1186/s12911-026-03814-5
Primary Topic
Interstitial Lung Diseases and Idiopathic Pulmonary Fibrosis
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Development and validation of an interpretable machine learning model for predicting in-hospital mortality in patients with pulmonary fibrosis

Yadong Yuan, Pei Wang, Xiaodan Jiao, Yanping Zhang
BMC Medical Informatics and Decision Making
Interstitial Lung Diseases and Idiopathic Pulmonary Fibrosis
article

Development and validation of an interpretable machine learning model for predicting in-hospital mortality in patients with pulmonary fibrosis

Yadong Yuan, Pei Wang, Xiaodan Jiao, Yanping Zhang
article en

Abstract

In-hospital mortality risk varies substantially among patients with pulmonary fibrosis (PF). Existing prognostic models have primarily focused on long-term survival and have limited ability to capture early in-hospital mortality risk or incorporate worsening oxygenation, inflammatory burden, and treatment-escalation information. This study aimed to develop and internally validate an interpretable machine learning model for in-hospital mortality risk stratification using early admission clinical data and information on ventilatory support within 24 h after admission in hospitalized patients with PF. This retrospective cohort study included hospitalized patients with PF admitted to a tertiary hospital between January 2019 and December 2023. Candidate variables included demographic characteristics, comorbidities, disease subtype, clinical symptoms, laboratory tests, arterial blood gas analysis, modified Medical Research Council (mMRC) dyspnea score, and the use of invasive or noninvasive ventilation within 24 h after admission. Nested LASSO feature selection combined with repeated 10-fold cross-validation was used for model development and internal validation. The main candidate models included logistic regression (LR), naive Bayes (NB), support vector machine (SVM), random forest (RF), neural network (NN), multilayer perceptron (MLP), adaptive boosting (AdaBoost), and gradient boosting machine (GBM). A supplementary clinically simplified model based on LASSO prescreening, recursive feature elimination, and logistic regression (RFE_LR) was also developed. Model performance was evaluated using the area under the receiver operating characteristic curve (AUC), threshold-dependent classification metrics, calibration analysis, decision curve analysis, clinical impact curves, bootstrap internal validation, and SHAP interpretation. Additional analyses included pairwise AUC comparisons with Holm adjustment for multiple testing, alternative-threshold analyses, exploratory in-sample recalibration analyses, and sensitivity analyses related to noninvasive ventilation (NIV). A total of 547 patients were included, of whom 108 met the prespecified primary mortality outcome, corresponding to an event rate of 19.74%. Six variables met the prespecified ≥ 50% selection-frequency criterion across the outer cross-validation folds: neutrophil percentage (NE%), NIV use within 24 h after admission, hypertension, mMRC score, idiopathic pulmonary fibrosis (IPF), and arterial oxygen partial pressure (PaO₂). Among the main candidate models, the NN model achieved the highest patient-level aggregated repeated-cross-validation AUC of 0.821 (nominal 95% CI: 0.772–0.869) and was selected as the final main model. This estimate was calculated after averaging each patient’s out-of-fold predicted probabilities across the 10 cross-validation repeats. After Holm adjustment for eight pairwise comparisons, only the AUC difference between the NN and NB models remained statistically significant (Holm-adjusted P = 0.040), whereas all other pairwise differences were not statistically significant. RFE_LR achieved an AUC of 0.789 (95% CI: 0.734–0.844), suggesting that the simplified interpretable model retained complementary value. At the maximum-F1 threshold of 0.57, the final NN model had a sensitivity of 0.602 and a specificity of 0.897. Lower thresholds improved death detection but reduced specificity. Calibration analysis showed that the model reflected mortality risk gradients, although raw predicted probabilities were systematically overestimated. The NIV-only AUC was 0.626. In the clinically relevant NIV = 0 subgroup, the final NN model achieved an AUC of 0.769; however, at the primary threshold of 0.57, sensitivity was only 0.456 and specificity was 0.912. After NIV was excluded and the models were redeveloped, the highest AUC was 0.790. These findings indicate that the model retained discriminative information beyond NIV status; however, the low sensitivity at the primary threshold limits its use as a stand-alone early-warning rule in patients not receiving NIV. Bootstrap internal validation yielded an optimism-corrected AUC of 0.788 and a mean OOB AUC of 0.718, with a 2.5th–97.5th percentile range of 0.604–0.817. Across the principal internal-validation approaches, AUC point estimates ranged from 0.718 to 0.821, indicating moderate but method-dependent discrimination. SHAP analysis identified mMRC score, IPF, PaO₂, NE%, hypertension, and NIV use within 24 h after admission as the main contributing variables. This study developed and internally validated an interpretable machine learning model for early in-hospital mortality risk stratification in hospitalized patients with PF. The model showed moderate, method-dependent discrimination and may be more appropriate for risk ranking than for stand-alone death detection or direct absolute-risk estimation. Performance differences among most evaluated algorithms were limited, and raw predicted probabilities were systematically overestimated. In the NIV = 0 subgroup, the primary threshold showed high specificity but low sensitivity, further limiting its use as a stand-alone early-warning rule. External validation and appropriate recalibration are required before clinical implementation.

BMC Medical Informatics and Decision Making
Hebei Medical University (CN), Second Hospital of Hebei Medical University (CN)
Good health and well-being
Openalex Percentile: Top 11%
Interstitial Lung Diseases and Idiopathic Pulmonary Fibrosis
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.