Intelligent Vehicular Routing and Optimal Resource Allocation Using Deep Reinforcement Learning for Urban VANETs

The proliferation of connected and autonomous vehicles has intensified the demand for robust Vehicular Ad-hoc Networks (VANETs). A critical challenge in urban VANETs is the co-design of efficient routing protocols and optimal communication resource allocation under highly dynamic and resource-constrained conditions. Traditional routing algorithms often fail to adapt to rapid topological changes, while conventional resource allocation schemes are not cognizant of the specific data flow requirements of multi-hop routes. This paper proposes a novel integrated framework, Deep Reinforcement Learning-based Vehicular Routing with Optimal Resource Allocation (DRL-VRORL), to address this joint problem. We model the urban VANET environment as a Markov Decision Process (MDP) where intelligent agents on vehicles collaboratively learn routing and resource allocation policies. The core of DRL-VRORL is a Multi-Agent Deep Deterministic Policy Gradient (MADDPG) architecture, enhanced with a centralized critic for coordinated learning and a distributed actor for scalable execution. The framework simultaneously optimizes for end-to-end packet delivery delay, packet delivery ratio (PDR), and network throughput. For resource allocation, we formulate a convex optimization problem for power and sub-channel allocation in Vehicle-to-Infrastructure (V2I) and Vehicle-to-Vehicle (V2V) links, which is solved efficiently using the Lagrange duality method, guided by the DRL agent's routing decisions. Extensive simulations conducted in SUMO and NS-3 demonstrate that DRL-VRORL significantly outperforms state-of-the-art protocols like AODV, DSR, and Q-learning-based routing. Specifically, DRL-VRORL achieves up to a 32% higher PDR, 45% lower average end-to-end delay, and 28% better aggregate network throughput while maintaining superior resource utilization efficiency.

Authors

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-09-14
DOI
https://doi.org/10.5281/zenodo.22745712
Primary Topic
Vehicular Ad Hoc Networks (VANETs)
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Intelligent Vehicular Routing and Optimal Resource Allocation Using Deep Reinforcement Learning for Urban VANETs

K.Thamizhmaran
Zenodo (CERN European Organization for Nuclear Research)
Vehicular Ad Hoc Networks (VANETs)
article

Intelligent Vehicular Routing and Optimal Resource Allocation Using Deep Reinforcement Learning for Urban VANETs

K.Thamizhmaran
article en

Abstract

The proliferation of connected and autonomous vehicles has intensified the demand for robust Vehicular Ad-hoc Networks (VANETs). A critical challenge in urban VANETs is the co-design of efficient routing protocols and optimal communication resource allocation under highly dynamic and resource-constrained conditions. Traditional routing algorithms often fail to adapt to rapid topological changes, while conventional resource allocation schemes are not cognizant of the specific data flow requirements of multi-hop routes. This paper proposes a novel integrated framework, Deep Reinforcement Learning-based Vehicular Routing with Optimal Resource Allocation (DRL-VRORL), to address this joint problem. We model the urban VANET environment as a Markov Decision Process (MDP) where intelligent agents on vehicles collaboratively learn routing and resource allocation policies. The core of DRL-VRORL is a Multi-Agent Deep Deterministic Policy Gradient (MADDPG) architecture, enhanced with a centralized critic for coordinated learning and a distributed actor for scalable execution. The framework simultaneously optimizes for end-to-end packet delivery delay, packet delivery ratio (PDR), and network throughput. For resource allocation, we formulate a convex optimization problem for power and sub-channel allocation in Vehicle-to-Infrastructure (V2I) and Vehicle-to-Vehicle (V2V) links, which is solved efficiently using the Lagrange duality method, guided by the DRL agent's routing decisions. Extensive simulations conducted in SUMO and NS-3 demonstrate that DRL-VRORL significantly outperforms state-of-the-art protocols like AODV, DSR, and Q-learning-based routing. Specifically, DRL-VRORL achieves up to a 32% higher PDR, 45% lower average end-to-end delay, and 28% better aggregate network throughput while maintaining superior resource utilization efficiency.

Zenodo (CERN European Organization for Nuclear Research)
Industry, innovation and infrastructure
Openalex Percentile: Top 20%
Vehicular Ad Hoc Networks (VANETs)
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.