Hard-label black-box model extraction attacks against network intrusion detection systems via generative adversarial networks
Machine-learning-based network intrusion detection systems (ML-NIDS) are now widely deployed as security services, yet they remain vulnerable to model extraction. Prior work typically relies on the soft-label confidence scores of the victim, whereas ML-NIDS in production usually return only the Top-1 prediction. Network traffic is also tabular and severely class-imbalanced, so techniques developed for image models do not transfer directly. We show that diverse ML-NIDS can be cloned with high fidelity from hard-label queries combined with a small proxy set, without any internal information or confidence scores. To this end, we propose a framework that jointly trains a clone model and a sample generator, so that the generator produces informative queries and adapts to the characteristics of tabular network data. The framework further stabilizes training and preserves high extraction quality on rare attack classes, which are easily degraded under conventional optimization. We evaluate the attack under realistic black-box conditions on NSL-KDD, CIC-IDS2017, and CSE-CIC-IDS2018, including a cross-dataset distribution shift. Across five architecturally distinct ML-NIDS on CIC-IDS2017, it attains 82–99% Mean Per-Class Fidelity under a limited query budget, showing that restricting API outputs to Top-1 predictions alone is insufficient to protect deployed ML-NIDS against model extraction.
Authors
- Seungsoo Nam (ORCID: https://orcid.org/0000-0001-9948-6140)
- Daeseon Choi (ORCID: https://orcid.org/0000-0002-1438-0265)
- Donguk Min (ORCID: https://orcid.org/0009-0003-8714-8883)
Institutions
- Soongsil University (KR)
Publication Details
- Journal
- Journal of Information Security and Applications
- Published
- 2026-09-29
- DOI
- https://doi.org/10.1016/j.jisa.2026.104652
- Primary Topic
- Adversarial Robustness in Machine Learning
- Type
- article
- Field-Weighted Citation Impact
- 0.00