Stratified Bootstrap or Stratified Cross-Validation: How to Do Reliable Model Validation Under Class Imbalance

Model validation is essential for evaluating how reliably a machine learning model performs on unseen data, especially in imbalanced classification where performance estimates can be affected by an underrepresentation of the minority class. Cross-validation methods are widely used for this purpose. However, they do not always provide direct information about the quality of individual validation samples or the stability of model performance when the training data changes. This study proposes a reliability-controlled stratified bootstrap procedure for model validation in imbalanced classification. The method applies class-aware bootstrap sampling and constructs the test sample from observations not used for training. In addition, each bootstrap run is evaluated according to predefined reliability criteria related to minority-class representation, preservation of class proportions and train–test independence. Model performance is assessed using standard classification metrics together with minority-class measures, variability across bootstrap repetitions and the train–test performance gap. The proposed stratified bootstrap is evaluated across six imbalanced datasets and six classification models and is compared with holdout train/test splits, k-fold cross-validation, stratified k-fold cross-validation and stratified shuffle splitting. The results show that the proposed bootstrap produces performance metrics consistent with the benchmark resampling procedures while providing additional information about validation reliability and model stability. In addition, the parallel implementation of the proposed bootstrap reduces computational time in many of the experiments, although the computational advantage depends on the dataset and classifier. Therefore, the proposed bootstrap can complement established resampling procedures when reliability control, minority-class evaluation and stability under repeated changes in the training data are important.

Authors

Institutions

Publication Details

Journal
Machine Learning and Knowledge Extraction
Published
2026-10-09
DOI
https://doi.org/10.3390/make8100326
Primary Topic
Imbalanced Data Classification Techniques
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

Stratified Bootstrap or Stratified Cross-Validation: How to Do Reliable Model Validation Under Class Imbalance

Borislava Toleva
Machine Learning and Knowledge Extraction
Imbalanced Data Classification Techniques
article

Stratified Bootstrap or Stratified Cross-Validation: How to Do Reliable Model Validation Under Class Imbalance

Borislava Toleva
article en

Abstract

Model validation is essential for evaluating how reliably a machine learning model performs on unseen data, especially in imbalanced classification where performance estimates can be affected by an underrepresentation of the minority class. Cross-validation methods are widely used for this purpose. However, they do not always provide direct information about the quality of individual validation samples or the stability of model performance when the training data changes. This study proposes a reliability-controlled stratified bootstrap procedure for model validation in imbalanced classification. The method applies class-aware bootstrap sampling and constructs the test sample from observations not used for training. In addition, each bootstrap run is evaluated according to predefined reliability criteria related to minority-class representation, preservation of class proportions and train–test independence. Model performance is assessed using standard classification metrics together with minority-class measures, variability across bootstrap repetitions and the train–test performance gap. The proposed stratified bootstrap is evaluated across six imbalanced datasets and six classification models and is compared with holdout train/test splits, k-fold cross-validation, stratified k-fold cross-validation and stratified shuffle splitting. The results show that the proposed bootstrap produces performance metrics consistent with the benchmark resampling procedures while providing additional information about validation reliability and model stability. In addition, the parallel implementation of the proposed bootstrap reduces computational time in many of the experiments, although the computational advantage depends on the dataset and classifier. Therefore, the proposed bootstrap can complement established resampling procedures when reliability control, minority-class evaluation and stability under repeated changes in the training data are important.

Machine Learning and Knowledge ExtractionVol. 8(10)
Plovdiv University (BG)
Openalex Percentile: Top 12%
Imbalanced Data Classification Techniques
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.