A Systematic Analysis of Oversampling Intensity and Decision Thresholds in Imbalanced Medical Classification Under Leakage-Free Cross-Validation
Class imbalance in medical classification can substantially affect model performance, particularly when resampling intensity and decision threshold are considered separately. This study systematically investigates their joint effects using the Pima Indians Diabetes, Heart Failure, and Thoracic Surgery datasets. Five classifiers were evaluated with Baseline, SMOTE, and ADASYN under multiple oversampling intensities and decision thresholds ranging from 0.30 to 0.70 using leakage-free, repeated, stratified 10-fold cross-validation with five repetitions. The best mean F1-scores were 0.6890 for PIMA, 0.7684 for Heart Failure, and 0.3375 for Thoracic Surgery, with the preferred threshold and oversampling intensity differing across datasets. A mixed-effects analysis confirmed a significant interaction between oversampling intensity and decision threshold (χ2(16) = 141.77, p < 0.001). In contrast, the matched difference between SMOTE and ADASYN was negligible (ΔF1 = 0.0009, p = 0.663). Thresholds below 0.50 were particularly beneficial for PIMA and Thoracic Surgery, whereas Heart Failure performed best near 0.50. Oversampling also tended to increase Brier scores, indicating that improvements in threshold-specific F1-score did not necessarily correspond to better probability calibration. These findings demonstrate that dataset-specific joint evaluation of oversampling intensity and decision threshold is more informative than selecting a resampling method in isolation.
Authors
- Kuan-Chu Lu
- Ting-Wei Wu
Institutions
- Feng Chia University (TW)
- Shih Hsin University (TW)
Publication Details
- Journal
- Mathematics
- Published
- 2026-09-21
- DOI
- https://doi.org/10.3390/math14183421
- Primary Topic
- Imbalanced Data Classification Techniques
- Type
- article
- Field-Weighted Citation Impact
- 0.00