Training-Data Axes in Imitation Learning Cascade into Hybrid Reinforcement-Learning Fine-Tuning: A Leave-One-Out Bridge Study Under Matched and Cascade Evaluation

Hybrid imitation-learning-to-reinforcement-learning (IL→RL) driving stacks are typically evaluated against a single fixed IL prior, leaving open whether IL training-data quality determines downstream RL outcomes and whether hybrid actuator decoupling (IL steers, RL controls only speed) isolates the speed controller from IL deficiencies. We evaluate five trajectory-IL checkpoints from a fixed conditional imitation learning (CIL) architecture (one matched baseline corpus plus four equal-budget leave-one-axis-out (LOO) ablations: drop_map, drop_weather, drop_traffic, drop_perturbation) under a 2 × 2 grid: matched proximal-policy optimisation (PPO) retraining versus frozen-PPO cascade evaluation, crossed with 0NPC versus 20NPC traffic (non-player characters; n = 3 PPO seeds per arm, wired v2-progressive reward). Matched co-training yields 442.6-unit cross-seed mean-reward spread at 0NPC (drop_map high, drop_perturbation low): the IL training corpus alone changes how well the same PPO recipe can learn speed control. Kendall τ against prior LOO pure-tier distance-importance is mild at matched 0NPC (τ=+0.333), stronger under matched 20NPC (τ=+0.667), mildly negative under cascade at 0NPC (τ=−0.333), and null under cascade at 20NPC (τ=0.00), so absolute gaps cascade more reliably than axis order. Under 20NPC traffic, cascade buffers the weak arms: the largest cascade-matched gap is on drop_perturbation (Δ=+428.3), with drop_weather still large (Δ=+205.5), meaning a baseline-trained speed head is more tolerant of a deficient IL prior at inference than a speed head co-trained against that prior. Frozen-head decoupling is therefore more stable at inference than per-axis co-training once multi-agent stress activates failure modes absent from zero-traffic training.

Authors

Institutions

Publication Details

Journal
Applied Sciences
Published
2026-10-05
DOI
https://doi.org/10.3390/app16199868
Primary Topic
Reinforcement Learning in Robotics
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

Training-Data Axes in Imitation Learning Cascade into Hybrid Reinforcement-Learning Fine-Tuning: A Leave-One-Out Bridge Study Under Matched and Cascade Evaluation

Claudiu Radu Pozna, Laurențiu Carabulea
Applied Sciences
Reinforcement Learning in Robotics
article

Training-Data Axes in Imitation Learning Cascade into Hybrid Reinforcement-Learning Fine-Tuning: A Leave-One-Out Bridge Study Under Matched and Cascade Evaluation

Claudiu Radu Pozna, Laurențiu Carabulea
article en

Abstract

Hybrid imitation-learning-to-reinforcement-learning (IL→RL) driving stacks are typically evaluated against a single fixed IL prior, leaving open whether IL training-data quality determines downstream RL outcomes and whether hybrid actuator decoupling (IL steers, RL controls only speed) isolates the speed controller from IL deficiencies. We evaluate five trajectory-IL checkpoints from a fixed conditional imitation learning (CIL) architecture (one matched baseline corpus plus four equal-budget leave-one-axis-out (LOO) ablations: drop_map, drop_weather, drop_traffic, drop_perturbation) under a 2 × 2 grid: matched proximal-policy optimisation (PPO) retraining versus frozen-PPO cascade evaluation, crossed with 0NPC versus 20NPC traffic (non-player characters; n = 3 PPO seeds per arm, wired v2-progressive reward). Matched co-training yields 442.6-unit cross-seed mean-reward spread at 0NPC (drop_map high, drop_perturbation low): the IL training corpus alone changes how well the same PPO recipe can learn speed control. Kendall τ against prior LOO pure-tier distance-importance is mild at matched 0NPC (τ=+0.333), stronger under matched 20NPC (τ=+0.667), mildly negative under cascade at 0NPC (τ=−0.333), and null under cascade at 20NPC (τ=0.00), so absolute gaps cascade more reliably than axis order. Under 20NPC traffic, cascade buffers the weak arms: the largest cascade-matched gap is on drop_perturbation (Δ=+428.3), with drop_weather still large (Δ=+205.5), meaning a baseline-trained speed head is more tolerant of a deficient IL prior at inference than a speed head co-trained against that prior. Frozen-head decoupling is therefore more stable at inference than per-axis co-training once multi-agent stress activates failure modes absent from zero-traffic training.

Applied SciencesVol. 16(19)
Transylvania University of Brașov (RO)
Openalex Percentile: Top 10%
Reinforcement Learning in Robotics
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Training-Data Axes in Imitation Learning Cascade into Hybrid Reinforcement-Learning Fine-Tuning: A Leave-One-Out Bridge Study Under Matched and Cascade Evaluation — Claudiu Radu Pozna, Laurențiu Carabulea · Applied Sciences (2026) | TGRS Research Map | TGRS