A dynamic explainable AI framework for real-time intrusion detection and automated alert prioritization in security operation centers
Abstract Security Operation Center (SOC) analysts face 50–100 security alerts per hour, leading to cognitive fatigue and delayed incident response. Existing intrusion detection systems (IDS) suffer from high false-positive rates and opaque black-box architectures that erode analyst trust and slow triage decisions. This paper proposes the Dynamic Explainable Framework (DEF), which couples a machine-learning detection backbone with inference-time SHAP and LIME explanations and an explanation-aware alert prioritization layer. DEF is evaluated on two benchmarks under an identical protocol: a stratified 60,900-flow subset of CICIoT2023 (eight classes) and a deduplicated 60,889-flow subset of UNSW-NB15 (ten classes). The XGBoost detection backbone attains 94.05 ± 0.08% accuracy with 88.49 ± 0.13% macro F1 at a 1.29% false-positive rate on CICIoT2023, and 88.26 ± 0.07% accuracy with 68.22 ± 0.22% macro F1 at a 1.25% false-positive rate on UNSW-NB15; false-positive rate and AUC-ROC therefore transfer essentially unchanged across domains (0.99 and 0.98). Measured against the alerts the detector actually generates (true plus false positives, 2,849 on CICIoT2023 and 2,912 on UNSW-NB15) rather than against all inspected flows, the prioritization layer removes roughly half of the residual false positives at the operating point (91 of 184 on CICIoT2023) while raising macro precision by 2.8 points, at a bounded cost in recall. The residual recall cost concentrates on stealthy, low-signature attack classes, and a per-class threshold analysis bounds the safe operating range. SHAP explanations are generated at 1.6 ms per alert (fidelity $$\rho = 0.79$$ against permutation importance) and LIME at 144.4 ms (top-10 feature stability 0.62), with a combined detection-plus-attribution latency of 3.17 ms per alert and a sustained throughput of 315 alerts per second on commodity CPU hardware. Transformer-based temporal and graph-based topology-context modules were additionally implemented and evaluated; their fusion does not surpass the tabular backbone on either benchmark, a negative result that we report and analyze. Comparing the two datasets isolates its cause: roughly half (51%) of the fusion’s shortfall on CICIoT2023 is attributable to that benchmark’s omission of per-flow host and timestamp identifiers, which forces proxy sequence and graph construction. The framework is designed and profiled for real-time operation on benchmark traffic traces; validation in a live SOC deployment remains future work.
Authors
- Mohammed Alsuhaibani (ORCID: https://orcid.org/0000-0001-6567-6413)
- Muhammad Usama Nazir (ORCID: https://orcid.org/0009-0005-0625-5404)
- Asri Bin Ngadi
Institutions
- Qassim University (SA)
- University of Technology Malaysia (MY)
Publication Details
- Journal
- Scientific Reports
- Published
- 2026-09-24
- DOI
- https://doi.org/10.1038/s41598-026-72373-6
- Primary Topic
- Network Security and Intrusion Detection
- Type
- article
- Field-Weighted Citation Impact
- 0.00