Knowledge-guided reinforcement learning for robust vehicle platoon control under communication and actuator failures

Connected and automated vehicle (CAV) platoon control is a knowledge-intensive engineering task that requires integrating domain expertise with adaptive decision-making to achieve multi-objective optimization under real-world operational uncertainties. This paper presents a hierarchical knowledge-guided deep reinforcement learning (DRL) framework based on proximal policy optimization (PPO) for reliable and adaptive platoon coordination. A robust consensus controller with formal stability guarantees addresses practical challenges including communication delays, actuator saturation and partial failures, and external disturbances. A control-error-sensitive graph attention network (CEGAT) is developed to adaptively weight inter-vehicle information according to communication availability and tracking error magnitudes. Domain knowledge derived from stability analysis is systematically embedded into the learning process through progressive policy guidance and continuous reward shaping, while a fixed-certificate linear matrix inequality (LMI)-based screening mechanism ensures that learned control parameters satisfy formal stability conditions before deployment. The framework is validated on multiple naturalistic driving datasets from the OpenACC database. Quantitative evaluations show that knowledge guidance halves training time relative to non-guided learning, reduces tracking errors by up to 32% over state-of-the-art methods, and outperforms commercial adaptive cruise control (ACC) systems by over 57% in velocity tracking accuracy. Extensive heterogeneous traffic experiments further confirm robust performance at human-driven vehicle (HDV) proportions up to 62.5%, insensitivity to HDV placement, and scalability to formations of 16 vehicles, demonstrating the framework’s generality for autonomous vehicle coordination in realistic mixed traffic environments.

Authors

Institutions

Publication Details

Journal
Advanced Engineering Informatics
Published
2026-09-11
DOI
https://doi.org/10.1016/j.aei.2026.105220
Primary Topic
Traffic control and management
Type
article
Field-Weighted Citation Impact
0.00

Funders

Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Knowledge-guided reinforcement learning for robust vehicle platoon control under communication and actuator failures

Zekai Lv, Zhihe Xu, Junhong Xie, Jianzhong Chen
Advanced Engineering Informatics
Traffic control and management
article

Knowledge-guided reinforcement learning for robust vehicle platoon control under communication and actuator failures

Zekai Lv, Zhihe Xu, Junhong Xie, Jianzhong Chen
article en

Abstract

Connected and automated vehicle (CAV) platoon control is a knowledge-intensive engineering task that requires integrating domain expertise with adaptive decision-making to achieve multi-objective optimization under real-world operational uncertainties. This paper presents a hierarchical knowledge-guided deep reinforcement learning (DRL) framework based on proximal policy optimization (PPO) for reliable and adaptive platoon coordination. A robust consensus controller with formal stability guarantees addresses practical challenges including communication delays, actuator saturation and partial failures, and external disturbances. A control-error-sensitive graph attention network (CEGAT) is developed to adaptively weight inter-vehicle information according to communication availability and tracking error magnitudes. Domain knowledge derived from stability analysis is systematically embedded into the learning process through progressive policy guidance and continuous reward shaping, while a fixed-certificate linear matrix inequality (LMI)-based screening mechanism ensures that learned control parameters satisfy formal stability conditions before deployment. The framework is validated on multiple naturalistic driving datasets from the OpenACC database. Quantitative evaluations show that knowledge guidance halves training time relative to non-guided learning, reduces tracking errors by up to 32% over state-of-the-art methods, and outperforms commercial adaptive cruise control (ACC) systems by over 57% in velocity tracking accuracy. Extensive heterogeneous traffic experiments further confirm robust performance at human-driven vehicle (HDV) proportions up to 62.5%, insensitivity to HDV placement, and scalability to formations of 16 vehicles, demonstrating the framework’s generality for autonomous vehicle coordination in realistic mixed traffic environments.

Advanced Engineering InformaticsVol. 77
Northwestern Polytechnical University (CN), Shenzhen Polytechnic University (CN)
Major Projects of Guangdong Education Department for Foundation Research and Applied Research
Climate action
Openalex Percentile: Top 15%
Traffic control and management
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.