Intelligent security policy generation algorithm based on graph neural network and reinforcement learning for SD-WAN
With the evolution of enterprise networks to multi-site and cloud architectures, Software-Defined Wide Area Network (SD-WAN) is prone to the problems of insufficient global policy regulation and frequent policy adjustment under dynamic security threats, which in turn leads to increased delay, aggravated packet loss and unstable network operation. Existing research mostly focuses on traffic optimization or local security detection, lacking global topology security situation modeling for SD-WAN and intelligent policy generation mechanisms that balance security benefits and policy stability. To address these issues, this study proposes an intelligent security policy generation model for dynamic security threats, Security-aware Graph-based Dual Reinforcement Learning for SD-WAN (SGDRL-SDWAN). The study achieves joint representation of network topology and security risks through constructing a security-aware graph neural network. Additionally, it designs a two-stage reinforcement learning framework with collaborative optimization of policy generation and stability constraints, realizing adaptive generation and stable control of security policies. The simulation environment is built based on open network traffic and intrusion detection data for verification. The experimental results show that SGDRL-SDWAN can reduce the end-to-end delay to 52.34 ms and the packet loss rate to 0.82% in the dynamic attack scenario, and effectively alleviate the problem of link load concentration. Under the conditions of persistent attacks and load fluctuation, the number of policy updates and network performance jitter are significantly lower than those of the comparison model. This study aims to provide a deployable intelligent security policy generation method for enterprise network operators and SD-WAN management platforms. It supports automatic generation and continuous optimization of security policies in dynamic attack environments, reducing control oscillations and operational burdens caused by manual policy adjustments, thereby enhancing network security and operational stability while lowering operational complexity.
Authors
- Guoqiang Ren (ORCID: https://orcid.org/0009-0000-9027-0343)
- Guang Cheng
- Hao Xu
- Ying Xu
- Guifu Xiao
Institutions
- Southeast University (BD)
- Southeast University (CN)
Publication Details
- Journal
- Discover Internet of Things
- Published
- 2026-09-30
- DOI
- https://doi.org/10.1007/s43926-026-00487-4
- Primary Topic
- Software-Defined Networks and 5G
- Type
- article
- Field-Weighted Citation Impact
- 0.00