A dynamic explainable AI framework for real-time intrusion detection and automated alert prioritization in security operation centers

Abstract Security Operation Center (SOC) analysts face 50–100 security alerts per hour, leading to cognitive fatigue and delayed incident response. Existing intrusion detection systems (IDS) suffer from high false-positive rates and opaque black-box architectures that erode analyst trust and slow triage decisions. This paper proposes the Dynamic Explainable Framework (DEF), which couples a machine-learning detection backbone with inference-time SHAP and LIME explanations and an explanation-aware alert prioritization layer. DEF is evaluated on two benchmarks under an identical protocol: a stratified 60,900-flow subset of CICIoT2023 (eight classes) and a deduplicated 60,889-flow subset of UNSW-NB15 (ten classes). The XGBoost detection backbone attains 94.05 ± 0.08% accuracy with 88.49 ± 0.13% macro F1 at a 1.29% false-positive rate on CICIoT2023, and 88.26 ± 0.07% accuracy with 68.22 ± 0.22% macro F1 at a 1.25% false-positive rate on UNSW-NB15; false-positive rate and AUC-ROC therefore transfer essentially unchanged across domains (0.99 and 0.98). Measured against the alerts the detector actually generates (true plus false positives, 2,849 on CICIoT2023 and 2,912 on UNSW-NB15) rather than against all inspected flows, the prioritization layer removes roughly half of the residual false positives at the operating point (91 of 184 on CICIoT2023) while raising macro precision by 2.8 points, at a bounded cost in recall. The residual recall cost concentrates on stealthy, low-signature attack classes, and a per-class threshold analysis bounds the safe operating range. SHAP explanations are generated at 1.6 ms per alert (fidelity $$\rho = 0.79$$ against permutation importance) and LIME at 144.4 ms (top-10 feature stability 0.62), with a combined detection-plus-attribution latency of 3.17 ms per alert and a sustained throughput of 315 alerts per second on commodity CPU hardware. Transformer-based temporal and graph-based topology-context modules were additionally implemented and evaluated; their fusion does not surpass the tabular backbone on either benchmark, a negative result that we report and analyze. Comparing the two datasets isolates its cause: roughly half (51%) of the fusion’s shortfall on CICIoT2023 is attributable to that benchmark’s omission of per-flow host and timestamp identifiers, which forces proxy sequence and graph construction. The framework is designed and profiled for real-time operation on benchmark traffic traces; validation in a live SOC deployment remains future work.

Authors

Institutions

Publication Details

Journal
Scientific Reports
Published
2026-09-24
DOI
https://doi.org/10.1038/s41598-026-72373-6
Primary Topic
Network Security and Intrusion Detection
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

A dynamic explainable AI framework for real-time intrusion detection and automated alert prioritization in security operation centers

Mohammed Alsuhaibani, Muhammad Usama Nazir, Asri Bin Ngadi
Scientific Reports
Network Security and Intrusion Detection
article

A dynamic explainable AI framework for real-time intrusion detection and automated alert prioritization in security operation centers

Mohammed Alsuhaibani, Muhammad Usama Nazir, Asri Bin Ngadi
article en

Abstract

Abstract Security Operation Center (SOC) analysts face 50–100 security alerts per hour, leading to cognitive fatigue and delayed incident response. Existing intrusion detection systems (IDS) suffer from high false-positive rates and opaque black-box architectures that erode analyst trust and slow triage decisions. This paper proposes the Dynamic Explainable Framework (DEF), which couples a machine-learning detection backbone with inference-time SHAP and LIME explanations and an explanation-aware alert prioritization layer. DEF is evaluated on two benchmarks under an identical protocol: a stratified 60,900-flow subset of CICIoT2023 (eight classes) and a deduplicated 60,889-flow subset of UNSW-NB15 (ten classes). The XGBoost detection backbone attains 94.05 ± 0.08% accuracy with 88.49 ± 0.13% macro F1 at a 1.29% false-positive rate on CICIoT2023, and 88.26 ± 0.07% accuracy with 68.22 ± 0.22% macro F1 at a 1.25% false-positive rate on UNSW-NB15; false-positive rate and AUC-ROC therefore transfer essentially unchanged across domains (0.99 and 0.98). Measured against the alerts the detector actually generates (true plus false positives, 2,849 on CICIoT2023 and 2,912 on UNSW-NB15) rather than against all inspected flows, the prioritization layer removes roughly half of the residual false positives at the operating point (91 of 184 on CICIoT2023) while raising macro precision by 2.8 points, at a bounded cost in recall. The residual recall cost concentrates on stealthy, low-signature attack classes, and a per-class threshold analysis bounds the safe operating range. SHAP explanations are generated at 1.6 ms per alert (fidelity $$\rho = 0.79$$ against permutation importance) and LIME at 144.4 ms (top-10 feature stability 0.62), with a combined detection-plus-attribution latency of 3.17 ms per alert and a sustained throughput of 315 alerts per second on commodity CPU hardware. Transformer-based temporal and graph-based topology-context modules were additionally implemented and evaluated; their fusion does not surpass the tabular backbone on either benchmark, a negative result that we report and analyze. Comparing the two datasets isolates its cause: roughly half (51%) of the fusion’s shortfall on CICIoT2023 is attributable to that benchmark’s omission of per-flow host and timestamp identifiers, which forces proxy sequence and graph construction. The framework is designed and profiled for real-time operation on benchmark traffic traces; validation in a live SOC deployment remains future work.

Scientific Reports
Qassim University (SA), University of Technology Malaysia (MY)
Peace, Justice and strong institutions
Openalex Percentile: Top 9%
Network Security and Intrusion Detection
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.