DriftGuard-SLA: A Paired Multi-Fault Evaluation of Prediction-Triggered Self-Correction in Cloud-Native Service-Oriented Architectures
We evaluated whether predictive alerts improve recovery in DriftGuard-SLA, a bounded controller for the Kubernetes-based DeathStarBench Hotel Reservation application. A logistic model trained on 35 monitoring runs was frozen before evaluation in 350 runs comprising 70 matched blocks, five policies, six faults, and a no-fault condition. For new sustained service-level objective breaches within 10 s, monitoring-run predictions achieved a Matthews correlation coefficient of 0.285, tie-aware average precision of 0.174, precision of 0.725, and recall of 0.119. Compared with monitoring alone, the complete controller reduced violation time by 7.70 s (95% confidence interval, 0.00–21.40; Holm-adjusted p = 0.031). Compared with fixed rules sharing its action mapping, the reduction was inconclusive at 1.08 s (−1.07–5.18; Holm-adjusted p = 0.453). Excluding four pre-injection actions yielded a descriptive sensitivity estimate of 0.02 s (−1.21–1.40). Of 38 faulted run actions, 4 preceded injection, 14 occurred during an active breach, 14 preceded a new onset, 5 had neither relation during complete observation, and one was censored. These findings support the feasibility of bounded intervention and paired evaluation, but do not establish superior predictive triggering, cost savings, or generalization beyond this benchmark.
Authors
- Bader Khalid Alshemaimri (ORCID: https://orcid.org/0009-0004-8865-5438)
- Ahmed Ghoneim (ORCID: https://orcid.org/0000-0003-2076-8925)
- Abdulaziz Jarallah Alghadban (ORCID: https://orcid.org/0009-0005-0519-1824)
Institutions
- King Saud University (SA)
Publication Details
- Journal
- Computation
- Published
- 2026-09-28
- DOI
- https://doi.org/10.3390/computation14100229
- Primary Topic
- Software System Performance and Reliability
- Type
- article
- Field-Weighted Citation Impact
- 0.00