Leakage-aware and imbalance-robust gradient boosting for intrusion detection in cyber-physical healthcare systems under adversarial AI threats

Abstract Cyber-physical healthcare systems (CPHS) increasingly depend on machine-learning intrusion detection systems (IDS) to protect networked medical devices and patient-monitoring infrastructure. Yet the IDS itself introduces a new attack surface: it is exposed to AI-specific threats—data poisoning, adversarial evasion, and model hijacking (backdoors)—and its reported accuracy is frequently inflated by undetected label leakage and by class imbalance handled in ways that distort the operating point. We present a leakage-aware, imbalance-robust detection pipeline for the WUSTL-EHMS-2020 healthcare CPS dataset (16,318 flows; 35 network-flow and 8 biometric features; 12.5% attack prevalence) and we audit it against all three threat families. A depth-one decision-stump audit identifies and removes a source-MAC column that alone predicts the label with accuracy 1.0 (a textbook leak), restoring honest evaluation. Class imbalance is handled inside a gradient-boosted detector through cost-sensitive re-weighting () rather than aggressive synthetic balancing, which we show is the cause of the low attack-class precision otherwise observed. The tuned detector attains held-out accuracy 0.973, attack-class precision 0.87, recall 0.93, $$F_1=0.90$$ , and AUC 0.992, and is the statistically best model across a ten-classifier benchmark (Friedman $$\\chi ^2=86.55$$ , $$p\\approx 8\\times 10^{-15}$$ ; pairwise Wilcoxon $$p=0.002$$ , significant after Bonferroni correction for nine simultaneous comparisons). Threat auditing on this dataset reveals a layered risk profile: label-flipping degrades $$F_1$$ modestly, adversarial evasion peaks at 0.45 under a large-budget FGSM attack, and all backdoor triggers achieve $$100\\%$$ bypass at $$1\\%$$ poisoning while preserving clean-traffic utility—a silent compromise that clean-data monitoring cannot detect; full backdoor defences (activation clustering, trigger reverse-engineering) are identified as priority future work rather than implemented here. As a model-agnostic defensive layer we add a randomized-smoothing conformal-prediction abstainer ( $$K=100$$ smoothing draws, $$\\sigma =0.15$$ ) that holds its $$95\\%$$ coverage guarantee (empirical 0.955) and re-flags evaded flows the point classifier misses. We further show that a per-sample detector calibrated to a $$5\\%$$ false-alarm rate tracks true evasion, whereas distribution-level tests fire on any perturbation regardless of success. The findings are grounded in a single healthcare cyber-physical benchmark and should be interpreted with appropriate caution before generalising to other datasets or attack repertoires; the methodology, however, is dataset-agnostic and directly transferable.

Authors

Institutions

Publication Details

Journal
Scientific Reports
Published
2026-09-18
DOI
https://doi.org/10.1038/s41598-026-70077-5
Primary Topic
Adversarial Robustness in Machine Learning
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Leakage-aware and imbalance-robust gradient boosting for intrusion detection in cyber-physical healthcare systems under adversarial AI threats

Wali Ahmed, Md Azizul Rahaman, Parves Alam, Shakhawat Hossain Refat et al.
Scientific Reports
Adversarial Robustness in Machine Learning
article

Leakage-aware and imbalance-robust gradient boosting for intrusion detection in cyber-physical healthcare systems under adversarial AI threats

Wali Ahmed, Md Azizul Rahaman, Parves Alam, Shakhawat Hossain Refat, Nikosi Zuberi, Shaon Reza, Jubayer Uddin Shamim, Pappu Roy
article en

Abstract

Abstract Cyber-physical healthcare systems (CPHS) increasingly depend on machine-learning intrusion detection systems (IDS) to protect networked medical devices and patient-monitoring infrastructure. Yet the IDS itself introduces a new attack surface: it is exposed to AI-specific threats—data poisoning, adversarial evasion, and model hijacking (backdoors)—and its reported accuracy is frequently inflated by undetected label leakage and by class imbalance handled in ways that distort the operating point. We present a leakage-aware, imbalance-robust detection pipeline for the WUSTL-EHMS-2020 healthcare CPS dataset (16,318 flows; 35 network-flow and 8 biometric features; 12.5% attack prevalence) and we audit it against all three threat families. A depth-one decision-stump audit identifies and removes a source-MAC column that alone predicts the label with accuracy 1.0 (a textbook leak), restoring honest evaluation. Class imbalance is handled inside a gradient-boosted detector through cost-sensitive re-weighting () rather than aggressive synthetic balancing, which we show is the cause of the low attack-class precision otherwise observed. The tuned detector attains held-out accuracy 0.973, attack-class precision 0.87, recall 0.93, $$F_1=0.90$$ , and AUC 0.992, and is the statistically best model across a ten-classifier benchmark (Friedman $$\chi ^2=86.55$$ , $$p\approx 8\times 10^{-15}$$ ; pairwise Wilcoxon $$p=0.002$$ , significant after Bonferroni correction for nine simultaneous comparisons). Threat auditing on this dataset reveals a layered risk profile: label-flipping degrades $$F_1$$ modestly, adversarial evasion peaks at 0.45 under a large-budget FGSM attack, and all backdoor triggers achieve $$100\%$$ bypass at $$1\%$$ poisoning while preserving clean-traffic utility—a silent compromise that clean-data monitoring cannot detect; full backdoor defences (activation clustering, trigger reverse-engineering) are identified as priority future work rather than implemented here. As a model-agnostic defensive layer we add a randomized-smoothing conformal-prediction abstainer ( $$K=100$$ smoothing draws, $$\sigma =0.15$$ ) that holds its $$95\%$$ coverage guarantee (empirical 0.955) and re-flags evaded flows the point classifier misses. We further show that a per-sample detector calibrated to a $$5\%$$ false-alarm rate tracks true evasion, whereas distribution-level tests fire on any perturbation regardless of success. The findings are grounded in a single healthcare cyber-physical benchmark and should be interpreted with appropriate caution before generalising to other datasets or attack repertoires; the methodology, however, is dataset-agnostic and directly transferable.

Scientific Reports
University of Bridgeport (US), Touro College (US), University of Kinshasa (CD), Open University Malaysia (MY)
Industry, innovation and infrastructure
Openalex Percentile: Top 8%
Adversarial Robustness in Machine Learning
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.