FROA-Drive: Failure-Routed Offline Adaptation for Lightweight Vision–Language–Action Autonomous Driving
Vision–Language–Action (VLA) models have shown strong potential for end-to-end autonomous driving, yet their post-training commonly relies on expensive simulator interaction or global policy updates. For an already competent pretrained policy, targeted correction is substantially more economical than repeated simulator interaction or global policy updates. We propose FROA-Drive, a lightweight offline adaptation framework that treats post-training as selective local correction. Starting from the frozen MindDrive-IL 0.5B checkpoint, FROA-Drive constructs a compact policy-centered candidate family around the original Top-1 speed trajectory, learns candidate-relative utility with a 1.51-million-parameter reranker, and uses a validation-calibrated gate to intervene only when an alternative is sufficiently supported. The original meta-actions and path trajectory are preserved by construction, providing an exact identity fallback and preventing unconstrained policy drift. On Chat-B2D-plus-10K, FROA-Drive achieves 72.8% candidate-selection accuracy and 84.2% intervention precision with a 5.4% proxy-violation rate. On Bench2Drive, it reaches a Driving Score of 77.08 and a Success Rate of 52.36%, improving the MindDrive-IL baseline by 1.23 DS and 3.06 percentage points in SR while requiring only 1.51 million trainable parameters and no online rollout during adaptation. These results establish selective local reranking as an effective, computation-efficient post-training strategy for lightweight driving VLA models.
Authors
- Ao Xu (ORCID: https://orcid.org/0009-0001-1763-2069)
- Yunhan Xu
Institutions
- Tongji University (CN)
- Nanjing University of Information Science and Technology (CN)
Publication Details
- Journal
- Sensors
- Published
- 2026-09-20
- DOI
- https://doi.org/10.3390/s26185962
- Primary Topic
- Multimodal Machine Learning Applications
- Type
- article
- Field-Weighted Citation Impact
- 0.00