A hierarchical deep reinforcement learning and model predictive control framework for safe longitudinal control and vehicle-to-grid energy management in autonomous electric vehicles

Current methods for autonomous electric vehicle (AEV) control treat longitudinal motion and energy management as decoupled problems: rule-based controllers (RBC) lack adaptability to dynamic pricing, while pure deep reinforcement learning (DRL) methods cannot guarantee the hard safety constraints required for production deployment. This paper proposes a hierarchical framework that resolves both limitations simultaneously, unifying longitudinal control and vehicle-to-grid (V2G) power dispatch in a single architecture. At the upper level, a Soft Actor-Critic deep reinforcement learning (SAC-DRL) agent learns a joint policy over traction commands and V2G dispatch, guided by a multi-component reward balancing safety, comfort, energy efficiency, and grid revenue. At the lower level, a Model Predictive Controller (MPC) enforces hard physical and safety constraints at 10 Hz, providing formal guarantees the learned SAC policy—a neural network with no built-in notion of physical constraints—cannot give on its own. The framework is experimentally evaluated through high-fidelity co-simulation using the Simulation of Urban Mobility (SUMO) platform, over 200 test episodes under diverse dynamic driving conditions: traffic densities of 200–800 veh/h, initial state-of-charge (SoC) from 0.40–0.80, ambient temperatures of 0–45 °C, and variable time-of-use (TOU) pricing. Compared with the RBC baseline, the proposed framework achieves 0.139 kWh/km net energy consumption (25.7 % reduction), $0.063/trip V2G revenue (65.8 % improvement over RBC), zero safety violations, and 100 % constraint satisfaction rate across all test episodes. The control-loop latency of 11.5 ms mean satisfies the 100 ms hard real-time budget—the maximum control-cycle period mandated by SAE J2735 and ISO 26262—leaving 77.6 % headroom that enables onboard deployment on production AEV embedded processors and real-time adaptation to live traffic and pricing signals.

Authors

Institutions

Publication Details

Journal
Scientific Reports
Published
2026-09-08
DOI
https://doi.org/10.1038/s41598-026-67186-6
Primary Topic
Electric and Hybrid Vehicle Technologies
Type
article
Field-Weighted Citation Impact
0.00

Funders

Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

A hierarchical deep reinforcement learning and model predictive control framework for safe longitudinal control and vehicle-to-grid energy management in autonomous electric vehicles

Appalabathula Venkatesh, Tousif Khan Nizami, Subrahmanyam Tanala, Fareed Ahmad
Scientific Reports
Electric and Hybrid Vehicle Technologies
article

A hierarchical deep reinforcement learning and model predictive control framework for safe longitudinal control and vehicle-to-grid energy management in autonomous electric vehicles

Appalabathula Venkatesh, Tousif Khan Nizami, Subrahmanyam Tanala, Fareed Ahmad
article en

Abstract

Current methods for autonomous electric vehicle (AEV) control treat longitudinal motion and energy management as decoupled problems: rule-based controllers (RBC) lack adaptability to dynamic pricing, while pure deep reinforcement learning (DRL) methods cannot guarantee the hard safety constraints required for production deployment. This paper proposes a hierarchical framework that resolves both limitations simultaneously, unifying longitudinal control and vehicle-to-grid (V2G) power dispatch in a single architecture. At the upper level, a Soft Actor-Critic deep reinforcement learning (SAC-DRL) agent learns a joint policy over traction commands and V2G dispatch, guided by a multi-component reward balancing safety, comfort, energy efficiency, and grid revenue. At the lower level, a Model Predictive Controller (MPC) enforces hard physical and safety constraints at 10 Hz, providing formal guarantees the learned SAC policy—a neural network with no built-in notion of physical constraints—cannot give on its own. The framework is experimentally evaluated through high-fidelity co-simulation using the Simulation of Urban Mobility (SUMO) platform, over 200 test episodes under diverse dynamic driving conditions: traffic densities of 200–800 veh/h, initial state-of-charge (SoC) from 0.40–0.80, ambient temperatures of 0–45 °C, and variable time-of-use (TOU) pricing. Compared with the RBC baseline, the proposed framework achieves 0.139 kWh/km net energy consumption (25.7 % reduction), $0.063/trip V2G revenue (65.8 % improvement over RBC), zero safety violations, and 100 % constraint satisfaction rate across all test episodes. The control-loop latency of 11.5 ms mean satisfies the 100 ms hard real-time budget—the maximum control-cycle period mandated by SAE J2735 and ISO 26262—leaving 77.6 % headroom that enables onboard deployment on production AEV embedded processors and real-time adaptation to live traffic and pricing signals.

Scientific Reports
King Fahd University of Petroleum and Minerals (SA), Gujarat Technological University (IN), SRM University (IN), Indian Institute of Management Visakhapatnam (IN)
SRM Institute of Science and Technology
Affordable and clean energy
Openalex Percentile: Top 18%
Electric and Hybrid Vehicle Technologies
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.