A Leakage-Aware Evaluation of a Multi-Scale Attention Temporal Convolutional Network for Binary Network Intrusion Detection

Reliable evaluation is as important as model design in benchmark-based network intrusion detection. This study evaluates a Multi-Scale Attention Temporal Convolutional Network (MS-ATCN) for binary intrusion detection under a leakage-aware and reproducibility-oriented protocol on CIC-IDS2017, NSL-KDD, and UNSW-NB15. MS-ATCN combines multi-scale temporal convolution, channel and temporal attention, and class-weighted focal loss over windows of flow-level records. Because these components are established techniques, the study focuses on their integrated empirical behavior rather than proposing a new learning paradigm. The evaluation includes repeated-seed experiments, seed-aligned Wilcoxon signed-rank tests with Holm correction, component ablations, an exploratory single-run record-order check, a single-seed CIC-IDS2017 day-level holdout, an UNSW-NB15 identifier-restoration stress test, exploratory one-at-a-time hyperparameter perturbations, and computational-efficiency measurements. The repeated-run means show that MS-ATCN is not consistently the best neural model: MLP has higher mean accuracy and Macro-F1 on CIC-IDS2017, XGBoost and LightGBM have the strongest repeated-run results on NSL-KDD and UNSW-NB15, and CNN has higher mean accuracy and Macro-F1 than MS-ATCN on UNSW-NB15. No seed-aligned main-model or ablation comparison reached the Holm-adjusted 0.05 threshold. With five paired seeds, however, the exact two-sided Wilcoxon test cannot attain an unadjusted p-value below 0.0625, so adjusted nonsignificance does not establish equivalence. The exploratory diagnostics indicate sensitivity to split policy and record ordering but do not establish a general benefit from local temporal context. The findings support cautious multi-metric interpretation, strong tabular baselines, repeated-seed reporting, and explicit separation of model-only efficiency from end-to-end deployment performance.

Authors

Institutions

Publication Details

Journal
Mathematics
Published
2026-09-16
DOI
https://doi.org/10.3390/math14183357
Primary Topic
Network Security and Intrusion Detection
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

A Leakage-Aware Evaluation of a Multi-Scale Attention Temporal Convolutional Network for Binary Network Intrusion Detection

Yang Yu, Minna Gao, Mingmei Chen, Jinliang Yuan et al.
Mathematics
Network Security and Intrusion Detection
article

A Leakage-Aware Evaluation of a Multi-Scale Attention Temporal Convolutional Network for Binary Network Intrusion Detection

Yang Yu, Minna Gao, Mingmei Chen, Jinliang Yuan, Le Gao, Dan Gong
article en

Abstract

Reliable evaluation is as important as model design in benchmark-based network intrusion detection. This study evaluates a Multi-Scale Attention Temporal Convolutional Network (MS-ATCN) for binary intrusion detection under a leakage-aware and reproducibility-oriented protocol on CIC-IDS2017, NSL-KDD, and UNSW-NB15. MS-ATCN combines multi-scale temporal convolution, channel and temporal attention, and class-weighted focal loss over windows of flow-level records. Because these components are established techniques, the study focuses on their integrated empirical behavior rather than proposing a new learning paradigm. The evaluation includes repeated-seed experiments, seed-aligned Wilcoxon signed-rank tests with Holm correction, component ablations, an exploratory single-run record-order check, a single-seed CIC-IDS2017 day-level holdout, an UNSW-NB15 identifier-restoration stress test, exploratory one-at-a-time hyperparameter perturbations, and computational-efficiency measurements. The repeated-run means show that MS-ATCN is not consistently the best neural model: MLP has higher mean accuracy and Macro-F1 on CIC-IDS2017, XGBoost and LightGBM have the strongest repeated-run results on NSL-KDD and UNSW-NB15, and CNN has higher mean accuracy and Macro-F1 than MS-ATCN on UNSW-NB15. No seed-aligned main-model or ablation comparison reached the Holm-adjusted 0.05 threshold. With five paired seeds, however, the exact two-sided Wilcoxon test cannot attain an unadjusted p-value below 0.0625, so adjusted nonsignificance does not establish equivalence. The exploratory diagnostics indicate sensitivity to split policy and record ordering but do not establish a general benefit from local temporal context. The findings support cautious multi-metric interpretation, strong tabular baselines, repeated-seed reporting, and explicit separation of model-only efficiency from end-to-end deployment performance.

MathematicsVol. 14(18)
PLA Rocket Force University of Engineering (CN), Air Force Engineering University (CN), Chinese People's Armed Police Force Engineering University (CN)
Openalex Percentile: Top 8%
Network Security and Intrusion Detection
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.