DKI-GMDM: domain-knowledge integrated and imbalance-guided diffusion model for network intrusion detection

The deployment of learning-based Network Intrusion Detection Systems (NIDS) is severely hindered by the class imbalance commonly observed in network traffic, where malicious activities account for only a small fraction of the data. Although generative data augmentation provides a promising way to alleviate this problem, existing methods such as generative adversarial network (GAN)-based oversampling and standard diffusion-based tabular generators may fail to sufficiently capture minority-attack distributions or may generate samples that violate important network-domain constraints. To address these challenges, we propose the Domain-Knowledge Integrated and Imbalance-Guided Diffusion Model (DKI-GMDM), a diffusion-based augmentation framework designed for imbalanced NIDS data. Specifically, we introduce a domain-knowledge constraint (DKC) module that incorporates protocol-related validity rules and statistical dependencies into the denoising process, thereby improving the semantic validity of generated traffic samples under a defined constraint set. We further develop an imbalance-aware guidance mechanism that adaptively adjusts the sampling trajectory according to class frequency, encouraging the generation of more class-specific samples for minority attack categories. Extensive experiments on selected class subsets of the UNSW-NB15 and CICIDS2017 benchmarks show that DKI-GMDM improves downstream detection performance over the selected baselines under the evaluated subset-based settings. On the selected UNSW-NB15 subset, DKI-GMDM improves the F1-scores of the highly imbalanced Shellcode and Worms classes by 7.71 and 6.52 percentage points, respectively, compared with Modelling Tabular Data with Diffusion Models (TabDDPM). Overall, these findings demonstrate the effectiveness of DKI-GMDM in alleviating the impact of class imbalance in NIDS training under the evaluated benchmark settings.

Authors

Institutions

Publication Details

Journal
PeerJ Computer Science
Published
2026-10-05
DOI
https://doi.org/10.7717/peerj-cs.4122
Primary Topic
Network Security and Intrusion Detection
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

DKI-GMDM: domain-knowledge integrated and imbalance-guided diffusion model for network intrusion detection

Depin Peng, Yunqi Liu, Yonggan Zhang, Liang Li
PeerJ Computer Science
Network Security and Intrusion Detection
article

DKI-GMDM: domain-knowledge integrated and imbalance-guided diffusion model for network intrusion detection

Depin Peng, Yunqi Liu, Yonggan Zhang, Liang Li
article en

Abstract

The deployment of learning-based Network Intrusion Detection Systems (NIDS) is severely hindered by the class imbalance commonly observed in network traffic, where malicious activities account for only a small fraction of the data. Although generative data augmentation provides a promising way to alleviate this problem, existing methods such as generative adversarial network (GAN)-based oversampling and standard diffusion-based tabular generators may fail to sufficiently capture minority-attack distributions or may generate samples that violate important network-domain constraints. To address these challenges, we propose the Domain-Knowledge Integrated and Imbalance-Guided Diffusion Model (DKI-GMDM), a diffusion-based augmentation framework designed for imbalanced NIDS data. Specifically, we introduce a domain-knowledge constraint (DKC) module that incorporates protocol-related validity rules and statistical dependencies into the denoising process, thereby improving the semantic validity of generated traffic samples under a defined constraint set. We further develop an imbalance-aware guidance mechanism that adaptively adjusts the sampling trajectory according to class frequency, encouraging the generation of more class-specific samples for minority attack categories. Extensive experiments on selected class subsets of the UNSW-NB15 and CICIDS2017 benchmarks show that DKI-GMDM improves downstream detection performance over the selected baselines under the evaluated subset-based settings. On the selected UNSW-NB15 subset, DKI-GMDM improves the F1-scores of the highly imbalanced Shellcode and Worms classes by 7.71 and 6.52 percentage points, respectively, compared with Modelling Tabular Data with Diffusion Models (TabDDPM). Overall, these findings demonstrate the effectiveness of DKI-GMDM in alleviating the impact of class imbalance in NIDS training under the evaluated benchmark settings.

PeerJ Computer ScienceVol. 12
Shaoxing University (CN), Wuhan University (CN), Zhejiang Institute of Communications (CN), Zhejiang Lab (CN)
Openalex Percentile: Top 9%
Network Security and Intrusion Detection
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.