SHAP-based feature relevance score technique for robust and explainable network anomaly classification
Abstract The accurate detection of anomalies in Border Gateway Protocol (BGP) traffic remains a critical challenge for ensuring the stability and security of global Internet routing. Conventional approaches often rely on overall accuracy or traditional feature selection techniques, which provide limited interpretability and fail to address the presence of redundant or noisy attributes in complex datasets. This study introduces a novel explainability-driven technique for BGP anomaly classification based on SHAP (SHapley Additive exPlanations). The proposed Feature Relevance Score (FRS) represents the key contribution, integrating global importance, class-specific impact, and directional stability into a unified metric for feature ranking. Each of these signals, taken alone, leaves features with distinct discriminative behaviour indistinguishable in the final ranking, and FRS is designed to recover this separation. This innovation not only enhances transparency but also enables more effective optimization of machine learning models. The experimental validation using the LSTM classifier demonstrates significant performance gains: targeted removal of low-impact features, particularly IPv4 prefix lengths, leads to a 10% improvement in recall for the outage anomaly class and noticeable gains in F1-score, while maintaining robustness across other classes. In contrast, removal of highly ranked features confirms the sensitivity and reliability of the proposed metric. Stability analysis across 20 independent random subsamples confirms that FRS produces consistent rankings, achieving a mean Spearman rank correlation of 0.828 (minimum 0.508), which exceeds global importance alone (0.752, minimum 0.304), class-specific scoring (0.813), and their combination without the directional component (0.805). The coefficient of variation remains below 7% across all features, and Marginal Rank Probability reaches 0.90 and 0.85 for representative features, compared to 0.40 and 0.20 for the class-specific baseline. The reported 10% recall gain for outage anomalies is further corroborated by a 50-run evaluation, where outage recall increases from 0.70 to 0.81 after removal of low-ranked prefix features, confirming that the improvement reflects a systematic tendency rather than a single-run artifact. The results show that combining interpretability with feature optimisation improves detection of underrepresented anomaly classes on the considered BGP benchmark, and support the use of explainability-driven feature ranking as a concrete decision-making tool in the feature-selection stage of model design.
Authors
- Мар’ян Кирик
- Кrzysztof Przystupa (ORCID: https://orcid.org/0000-0003-4361-2763)
- Jaroslaw Sikora (ORCID: https://orcid.org/0000-0003-4843-0731)
- Stanislav Maruniak (ORCID: https://orcid.org/0009-0006-0635-512X)
- Volodymyr Rykhva (ORCID: https://orcid.org/0009-0008-2711-547X)
- Mykola Beshley
- Orest Kochan
- Halyna Beshley
Publication Details
- Journal
- Scientific Reports
- Published
- 2026-09-29
- DOI
- https://doi.org/10.1038/s41598-026-71731-8
- Primary Topic
- Network Security and Intrusion Detection
- Type
- article
- Field-Weighted Citation Impact
- 0.00