Development of a prediction model for pre-scan diagnostic categories in bone scintigraphy using multinomial logistic regression and XGBoost

We developed and compared three-class diagnostic prediction models based on multinomial logistic regression and XGBoost. By predicting diagnostic outcomes, the models are intended to provide an adjunctive reference for risk stratification in patients scheduled for whole-body bone scintigraphy. After excluding cases with missing values, 3,958 patients were retained from an initial cohort of 3,969. The entire dataset was chronologically divided into a temporal training set and a temporal validation set at an 8:2 ratio according to the date of examination. The temporal training set was then further randomly partitioned into a random training set and a random test set (also 8:2). Models were built using 11 clinical features. Hyperparameters were optimized on the random training set via grid search combined with 5-fold cross-validation. Using the selected optimal parameters, the models were retrained on both the random training set and the temporal training set, and their final performance was evaluated on the random test set and the temporal validation set, respectively. Evaluation metrics included accuracy, F1 score, area under the receiver operating characteristic curve (AUC), Brier score, calibration slope, and net benefit derived from decision curve analysis. To account for potential bias due to temporal drift, post-hoc calibration—Platt scaling for logistic regression and isotonic regression for XGBoost—was fitted via 5-fold cross-validation on the temporal training set. The fitted calibrators were then applied independently to the temporal validation set. Calibration performance was compared before and after adjustment to quantify the impact of temporal shift on the reliability of predicted probabilities. Both models demonstrated encouraging three-class diagnostic prediction performance on the random test set, with no statistically significant differences between the two models across all evaluated metrics (all p > 0.05). History of cancer and history of trauma emerged as the most important predictors. The models showed acceptable calibration performance. Decision curve analysis provided exploratory evidence of potential net benefit, but this potential benefit diminished in temporal validation. In the temporal validation set, pre-calibration performance declined to some extent, whereas post-calibration probability reliability improved. Both models effectively predicted diagnostic categories in bone scintigraphy, providing an adjunctive reference for pre-scan clinical risk assessment in patients scheduled for whole-body bone scintigraphy. Temporal validation showed decreased performance and diminished exploratory net benefit for the uncalibrated models, whereas probability calibration improved the reliability of predicted probabilities in this temporal cohort. These findings suggest that calibration may be useful for improving probability reliability, but they do not establish temporal generalizability, discrimination improvement, or readiness for clinical deployment; external validation is required.

Authors

Institutions

Publication Details

Journal
BMC Medical Imaging
Published
2026-10-06
DOI
https://doi.org/10.1186/s12880-026-02848-5
Primary Topic
Artificial Intelligence in Healthcare
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

Development of a prediction model for pre-scan diagnostic categories in bone scintigraphy using multinomial logistic regression and XGBoost

Peilin Wu, Ziyang Du, Kuo Ma
BMC Medical Imaging
Artificial Intelligence in Healthcare
article

Development of a prediction model for pre-scan diagnostic categories in bone scintigraphy using multinomial logistic regression and XGBoost

Peilin Wu, Ziyang Du, Kuo Ma
article en

Abstract

We developed and compared three-class diagnostic prediction models based on multinomial logistic regression and XGBoost. By predicting diagnostic outcomes, the models are intended to provide an adjunctive reference for risk stratification in patients scheduled for whole-body bone scintigraphy. After excluding cases with missing values, 3,958 patients were retained from an initial cohort of 3,969. The entire dataset was chronologically divided into a temporal training set and a temporal validation set at an 8:2 ratio according to the date of examination. The temporal training set was then further randomly partitioned into a random training set and a random test set (also 8:2). Models were built using 11 clinical features. Hyperparameters were optimized on the random training set via grid search combined with 5-fold cross-validation. Using the selected optimal parameters, the models were retrained on both the random training set and the temporal training set, and their final performance was evaluated on the random test set and the temporal validation set, respectively. Evaluation metrics included accuracy, F1 score, area under the receiver operating characteristic curve (AUC), Brier score, calibration slope, and net benefit derived from decision curve analysis. To account for potential bias due to temporal drift, post-hoc calibration—Platt scaling for logistic regression and isotonic regression for XGBoost—was fitted via 5-fold cross-validation on the temporal training set. The fitted calibrators were then applied independently to the temporal validation set. Calibration performance was compared before and after adjustment to quantify the impact of temporal shift on the reliability of predicted probabilities. Both models demonstrated encouraging three-class diagnostic prediction performance on the random test set, with no statistically significant differences between the two models across all evaluated metrics (all p > 0.05). History of cancer and history of trauma emerged as the most important predictors. The models showed acceptable calibration performance. Decision curve analysis provided exploratory evidence of potential net benefit, but this potential benefit diminished in temporal validation. In the temporal validation set, pre-calibration performance declined to some extent, whereas post-calibration probability reliability improved. Both models effectively predicted diagnostic categories in bone scintigraphy, providing an adjunctive reference for pre-scan clinical risk assessment in patients scheduled for whole-body bone scintigraphy. Temporal validation showed decreased performance and diminished exploratory net benefit for the uncalibrated models, whereas probability calibration improved the reliability of predicted probabilities in this temporal cohort. These findings suggest that calibration may be useful for improving probability reliability, but they do not establish temporal generalizability, discrimination improvement, or readiness for clinical deployment; external validation is required.

BMC Medical Imaging
Dongzhimen Hospital Affiliated to Beijing University of Chinese Medicine (CN)
Openalex Percentile: Top 7%
Artificial Intelligence in Healthcare
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.