HERA-GM: Evaluating Conditional Execution Authority for Offline Reinforcement Learning in Tactical Driving
HERA-GM separates an offline learned tactical proposal from the authority to execute it. A Dueling Double DQN trained with Conservative Q-Learning proposes one of five tactical actions. The runtime process then uses the action-value margin, semantic feature availability, annotation density, a Mahalanobis diagnostic, and fixed hard-rule conditions to assign ACCEPT, DEFER, or RECOVER. The study separately examined proposer agreement, authority changes, closed-loop outcomes, and held-out-family discrimination. The original frozen evaluation used 113 nuPlan Mini scenarios from 39 logs and 1970 closed-loop runs. An additional exploratory behavior-cloning block added 339 runs. Behavior cloning had slightly higher offline macro-F1 than CQL, whereas DDQN without CQL had much lower agreement under the tested configurations. The main comparison between M1 and the simpler B3 gate showed no supported primary safety difference, indicating limited added endpoint effect from Mahalanobis and hard-rule evidence in this cohort. M1 also showed lower safety-failure and drivable-area violation rates than behavior cloning, but with lower conditional progress; the primary result did not remain below 0.05 after pooled Holm adjustment across the five clean comparisons. The Mahalanobis score did not distinguish held-out semantic families reliably. The findings describe the operating trade-offs and limits of conditional execution authority rather than a safety guarantee.
Authors
- Mohammad Al Khaldy (ORCID: https://orcid.org/0009-0009-7502-4668)
- Ameen Shaheen (ORCID: https://orcid.org/0009-0003-3120-1892)
- Youcef Gheraibia (ORCID: https://orcid.org/0000-0002-0854-5211)
Institutions
- Al-Zaytoonah University of Jordan (JO)
- Petra University (JO)
- De Montfort University (GB)
Publication Details
- Journal
- Computation
- Published
- 2026-09-10
- DOI
- https://doi.org/10.3390/computation14090213
- Primary Topic
- Adversarial Robustness in Machine Learning
- Type
- article
- Field-Weighted Citation Impact
- 0.00