Delay-aware multi-agent reinforcement learning for mixed-autonomy platoon control

It is recognized that controlling mixed-autonomy platoons comprising connected and automated vehicles (CAVs) and human-driven vehicles (HDVs) can enhance traffic flow. While multi-agent reinforcement learning (MARL) is a promising real-time control paradigm, existing MARL-based platoon controllers rarely account for communication and execution delays, despite their inevitability in practice and their critical impact on safety and stability. Moreover, most delay-compensation approaches are model-based, which becomes unsuitable when HDV dynamics are unknown and model-free coordination among CAVs is required. To address this gap, we formulate mixed-autonomy platoon control with delays as a delayed Markov game and develop a delay-aware learning framework supported by a delay-dependent theoretical analysis. Specifically, we provide a theoretical analysis that establishes explicit performance bounds between delayed and undelayed tasks under smoothness conditions, highlighting that the delay-induced performance gap is governed by the policy smoothness and the belief uncertainty under delayed observations. Motivated by this insight, we propose a multi-agent transformer (MAT) that exploits the disturbance-propagation structure of platoons to learn coordinated and regularized policies, serving as effective undelayed experts for reliable transfer to delayed environments. Finally, we validate the proposed approach through extensive simulations and human-in-the-loop experiments, demonstrating the control performance and sample efficiency of the proposed method.

Authors

Institutions

Publication Details

Journal
Transportation Research Part C Emerging Technologies
Published
2026-10-07
DOI
https://doi.org/10.1016/j.trc.2026.106054
Primary Topic
Traffic control and management
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

Delay-aware multi-agent reinforcement learning for mixed-autonomy platoon control

Jingyuan Zhou, Kaidi Yang, Zhicheng Wang
Transportation Research Part C Emerging Technologies
Traffic control and management
article

Delay-aware multi-agent reinforcement learning for mixed-autonomy platoon control

Jingyuan Zhou, Kaidi Yang, Zhicheng Wang
article en

Abstract

It is recognized that controlling mixed-autonomy platoons comprising connected and automated vehicles (CAVs) and human-driven vehicles (HDVs) can enhance traffic flow. While multi-agent reinforcement learning (MARL) is a promising real-time control paradigm, existing MARL-based platoon controllers rarely account for communication and execution delays, despite their inevitability in practice and their critical impact on safety and stability. Moreover, most delay-compensation approaches are model-based, which becomes unsuitable when HDV dynamics are unknown and model-free coordination among CAVs is required. To address this gap, we formulate mixed-autonomy platoon control with delays as a delayed Markov game and develop a delay-aware learning framework supported by a delay-dependent theoretical analysis. Specifically, we provide a theoretical analysis that establishes explicit performance bounds between delayed and undelayed tasks under smoothness conditions, highlighting that the delay-induced performance gap is governed by the policy smoothness and the belief uncertainty under delayed observations. Motivated by this insight, we propose a multi-agent transformer (MAT) that exploits the disturbance-propagation structure of platoons to learn coordinated and regularized policies, serving as effective undelayed experts for reliable transfer to delayed environments. Finally, we validate the proposed approach through extensive simulations and human-in-the-loop experiments, demonstrating the control performance and sample efficiency of the proposed method.

Transportation Research Part C Emerging TechnologiesVol. 194
National University of Singapore (SG)
Openalex Percentile: Top 16%
Traffic control and management
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Delay-aware multi-agent reinforcement learning for mixed-autonomy platoon control — Jingyuan Zhou, Kaidi Yang, et al. · Transportation Research Part C Emerging Technologies (2026) | TGRS Research Map | TGRS