Development of a prediction model for pre-scan diagnostic categories in bone scintigraphy using multinomial logistic regression and XGBoost
We developed and compared three-class diagnostic prediction models based on multinomial logistic regression and XGBoost. By predicting diagnostic outcomes, the models are intended to provide an adjunctive reference for risk stratification in patients scheduled for whole-body bone scintigraphy. After excluding cases with missing values, 3,958 patients were retained from an initial cohort of 3,969. The entire dataset was chronologically divided into a temporal training set and a temporal validation set at an 8:2 ratio according to the date of examination. The temporal training set was then further randomly partitioned into a random training set and a random test set (also 8:2). Models were built using 11 clinical features. Hyperparameters were optimized on the random training set via grid search combined with 5-fold cross-validation. Using the selected optimal parameters, the models were retrained on both the random training set and the temporal training set, and their final performance was evaluated on the random test set and the temporal validation set, respectively. Evaluation metrics included accuracy, F1 score, area under the receiver operating characteristic curve (AUC), Brier score, calibration slope, and net benefit derived from decision curve analysis. To account for potential bias due to temporal drift, post-hoc calibration—Platt scaling for logistic regression and isotonic regression for XGBoost—was fitted via 5-fold cross-validation on the temporal training set. The fitted calibrators were then applied independently to the temporal validation set. Calibration performance was compared before and after adjustment to quantify the impact of temporal shift on the reliability of predicted probabilities. Both models demonstrated encouraging three-class diagnostic prediction performance on the random test set, with no statistically significant differences between the two models across all evaluated metrics (all p > 0.05). History of cancer and history of trauma emerged as the most important predictors. The models showed acceptable calibration performance. Decision curve analysis provided exploratory evidence of potential net benefit, but this potential benefit diminished in temporal validation. In the temporal validation set, pre-calibration performance declined to some extent, whereas post-calibration probability reliability improved. Both models effectively predicted diagnostic categories in bone scintigraphy, providing an adjunctive reference for pre-scan clinical risk assessment in patients scheduled for whole-body bone scintigraphy. Temporal validation showed decreased performance and diminished exploratory net benefit for the uncalibrated models, whereas probability calibration improved the reliability of predicted probabilities in this temporal cohort. These findings suggest that calibration may be useful for improving probability reliability, but they do not establish temporal generalizability, discrimination improvement, or readiness for clinical deployment; external validation is required.
Authors
- Peilin Wu
- Ziyang Du
- Kuo Ma
Institutions
- Dongzhimen Hospital Affiliated to Beijing University of Chinese Medicine (CN)
Publication Details
- Journal
- BMC Medical Imaging
- Published
- 2026-10-06
- DOI
- https://doi.org/10.1186/s12880-026-02848-5
- Primary Topic
- Artificial Intelligence in Healthcare
- Type
- article
- Field-Weighted Citation Impact
- 0.00