Deep reinforcement learning-optimized adaptive authenticated encryption and dynamic transmission control for secure LoRaWAN IoT communication

The growth of Internet of Things (IoT) devices in the industrial, environmental and urban sectors has raised the need for secure, energy efficient, and adaptive communication solutions. The LoRa (Long Range) technology combined with the LoRaWAN protocol stack is very promising due to its low power consumption, ability to operate far away, and scalability in resource-constrained deployment. The static encryption keys and fixed transmission parameters of conventional LoRaWAN implementations are, however, extremely vulnerable to eavesdropping, replay attacks, key compromise and brute force attacks. At the same time, the spreading factor, transmission power and coding rates are not matched to the dynamic channel characteristics of the wireless channel, which results in high consumption of energy, high packet loss rate and poor quality of service. This paper introduces a Deep Reinforcement Learning (DRL)-based framework that combines the adaptive use of Authenticated Encryption with Associated Data (AEAD) and the dynamic optimization of transmission with LoRaWAN-based IoT systems. The proposed architecture utilizes a Dueling Deep Q-Network (DDQN) agent which optimizes the entropy parameters for key generation and LoRa transmission parameters from real-time observations of the wireless channel, accompanied by security threat indices. Dynamic encryption keys are built from hardware sourced entropy (ADC thermal noise, clock-cycle jitter) and monotonic timestamps, resulting in key matrices which are unique for each session, and have measured entropy of greater than 7.95 bits per byte. A plaintext randomization pre-processing layer is developed based on eigenvectors to enhance the plaintext randomization before authenticated AES-GCM encryption, which offers confidentiality, integrity and authenticity guarantees. The problem of optimising spreading factor (SF), transmission power (TP), and coding rate (CR) with the aim of maximising packet delivery ratio, minimising energy consumption, latency and key entropy is formalised using a multi-objective reward function through a Markov Decision Process (MDP). The hardware validation is conducted on 20 real-time transmission cycles over a controlled 15-metre line-of-sight range, with 500 training episodes; results beyond this range are analytical projections rather than direct hardware measurements, as detailed in Sect. 6.1, and the results are: 93.7–96.5% mean data-recovery rate across the 20 transmission cycles (the percentage of transmitted sensor readings correctly decrypted and matched at the receiver), packet delivery ratio was improved by 27.8%, energy consumption was reduced by 32.5%, and the measured replay acceptance rate was zero across all tested replay scenarios (Table 2). The proposed integrated approach outperforms the other methods, including static LoRaWAN, AES-128 pre-shared key, CRC-only verification, and RL-only transmission optimization, in terms of all performance metrics evaluated on this single-node, 20-cycle, 15-metre hardware trial; broader statistical validation across repeated trials and multiple deployment scenarios is identified as future work (Sect. 8.1).

Authors

Institutions

Publication Details

Journal
Discover Internet of Things
Published
2026-09-25
DOI
https://doi.org/10.1007/s43926-026-00515-3
Primary Topic
IoT Networks and Protocols
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Deep reinforcement learning-optimized adaptive authenticated encryption and dynamic transmission control for secure LoRaWAN IoT communication

S. Vasudevan, P. Sundaravadivel, G. Abinaya, K. Muthulakshmi
Discover Internet of Things
IoT Networks and Protocols
article

Deep reinforcement learning-optimized adaptive authenticated encryption and dynamic transmission control for secure LoRaWAN IoT communication

S. Vasudevan, P. Sundaravadivel, G. Abinaya, K. Muthulakshmi
article en

Abstract

The growth of Internet of Things (IoT) devices in the industrial, environmental and urban sectors has raised the need for secure, energy efficient, and adaptive communication solutions. The LoRa (Long Range) technology combined with the LoRaWAN protocol stack is very promising due to its low power consumption, ability to operate far away, and scalability in resource-constrained deployment. The static encryption keys and fixed transmission parameters of conventional LoRaWAN implementations are, however, extremely vulnerable to eavesdropping, replay attacks, key compromise and brute force attacks. At the same time, the spreading factor, transmission power and coding rates are not matched to the dynamic channel characteristics of the wireless channel, which results in high consumption of energy, high packet loss rate and poor quality of service. This paper introduces a Deep Reinforcement Learning (DRL)-based framework that combines the adaptive use of Authenticated Encryption with Associated Data (AEAD) and the dynamic optimization of transmission with LoRaWAN-based IoT systems. The proposed architecture utilizes a Dueling Deep Q-Network (DDQN) agent which optimizes the entropy parameters for key generation and LoRa transmission parameters from real-time observations of the wireless channel, accompanied by security threat indices. Dynamic encryption keys are built from hardware sourced entropy (ADC thermal noise, clock-cycle jitter) and monotonic timestamps, resulting in key matrices which are unique for each session, and have measured entropy of greater than 7.95 bits per byte. A plaintext randomization pre-processing layer is developed based on eigenvectors to enhance the plaintext randomization before authenticated AES-GCM encryption, which offers confidentiality, integrity and authenticity guarantees. The problem of optimising spreading factor (SF), transmission power (TP), and coding rate (CR) with the aim of maximising packet delivery ratio, minimising energy consumption, latency and key entropy is formalised using a multi-objective reward function through a Markov Decision Process (MDP). The hardware validation is conducted on 20 real-time transmission cycles over a controlled 15-metre line-of-sight range, with 500 training episodes; results beyond this range are analytical projections rather than direct hardware measurements, as detailed in Sect. 6.1, and the results are: 93.7–96.5% mean data-recovery rate across the 20 transmission cycles (the percentage of transmitted sensor readings correctly decrypted and matched at the receiver), packet delivery ratio was improved by 27.8%, energy consumption was reduced by 32.5%, and the measured replay acceptance rate was zero across all tested replay scenarios (Table 2). The proposed integrated approach outperforms the other methods, including static LoRaWAN, AES-128 pre-shared key, CRC-only verification, and RL-only transmission optimization, in terms of all performance metrics evaluated on this single-node, 20-cycle, 15-metre hardware trial; broader statistical validation across repeated trials and multiple deployment scenarios is identified as future work (Sect. 8.1).

Discover Internet of ThingsVol. 6(1)
Vel Tech Rangarajan Dr. Sagunthala R&D Institute of Science and Technology (IN), Saveetha University (IN)
Affordable and clean energy
Openalex Percentile: Top 21%
IoT Networks and Protocols
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.