Adaptive Reinforcement Learning for Reliable and Sustainable Intelligent Vehicular Communication

Vehicular ad hoc networks (VANETs) require routing decisions that balance connectivity, link reliability and energy efficiency under continuously changing mobility and density. Conventional clustering-based routing combines these objectives using weights fixed in advance, which cannot reflect the changing relative importance of the objectives as conditions vary. This paper presents ISPY, a hybrid reinforcement learning framework for cluster-based routing in which the objective weighting is itself the learned quantity. A dueling double deep Q-network selects a discrete operating mode from a seven-dimensional macroscopic state comprising vehicle density, mean speed, speed variation, hop count, residual energy, energy efficiency and packet delivery ratio, while a Dirichlet-based continuous actor generates a normalised weight vector over anchor connectivity, link reliability and link persistence. Candidate next hops are ranked by a cost-based priority function and selected by minimisation. The reward combines packet delivery ratio, energy efficiency, and forwarding stability with coefficients 0.5, 0.3, and 0.2. The framework is evaluated on a vehicular mobility and communication trace segmented into 300 temporally aligned snapshots, compared against a fixed equal-weight policy that isolates the contribution of weight adaptation and against the conventional schemes N-HOP, VMaSC, and DMCNF under identical snapshots, transmission range, and link metric inputs. Training converges under early stopping at episode 130, with an average composite reward of 0.8161 against 0.8158 for the fixed-weight policy; this difference of 0.0003 is reported as a convergence signal. The learned weights vary systematically with network state. Comparison against learning-based baselines and a component-wise ablation are identified as necessary further work.

Authors

Institutions

Publication Details

Journal
Future Internet
Published
2026-09-25
DOI
https://doi.org/10.3390/fi18100509
Primary Topic
Vehicular Ad Hoc Networks (VANETs)
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Adaptive Reinforcement Learning for Reliable and Sustainable Intelligent Vehicular Communication

Zahoor Ur Rehman, Khalid Mahmood Awan, Shahid Kamal, Areeba Naseem Khan
Future Internet
Vehicular Ad Hoc Networks (VANETs)
article

Adaptive Reinforcement Learning for Reliable and Sustainable Intelligent Vehicular Communication

Zahoor Ur Rehman, Khalid Mahmood Awan, Shahid Kamal, Areeba Naseem Khan
article en

Abstract

Vehicular ad hoc networks (VANETs) require routing decisions that balance connectivity, link reliability and energy efficiency under continuously changing mobility and density. Conventional clustering-based routing combines these objectives using weights fixed in advance, which cannot reflect the changing relative importance of the objectives as conditions vary. This paper presents ISPY, a hybrid reinforcement learning framework for cluster-based routing in which the objective weighting is itself the learned quantity. A dueling double deep Q-network selects a discrete operating mode from a seven-dimensional macroscopic state comprising vehicle density, mean speed, speed variation, hop count, residual energy, energy efficiency and packet delivery ratio, while a Dirichlet-based continuous actor generates a normalised weight vector over anchor connectivity, link reliability and link persistence. Candidate next hops are ranked by a cost-based priority function and selected by minimisation. The reward combines packet delivery ratio, energy efficiency, and forwarding stability with coefficients 0.5, 0.3, and 0.2. The framework is evaluated on a vehicular mobility and communication trace segmented into 300 temporally aligned snapshots, compared against a fixed equal-weight policy that isolates the contribution of weight adaptation and against the conventional schemes N-HOP, VMaSC, and DMCNF under identical snapshots, transmission range, and link metric inputs. Training converges under early stopping at episode 130, with an average composite reward of 0.8161 against 0.8158 for the fixed-weight policy; this difference of 0.0003 is reported as a convergence signal. The learned weights vary systematically with network state. Comparison against learning-based baselines and a component-wise ablation are identified as necessary further work.

Future InternetVol. 18(10)
Sakarya University (TR), COMSATS University Islamabad (PK), Multimedia University (MY)
Affordable and clean energy
Openalex Percentile: Top 21%
Vehicular Ad Hoc Networks (VANETs)
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.