Application of the W-shaped process for a Reinforcement Learning use case on drone navigation
The interest in Machine Learning (ML) has spread to safety systems due to its ability to solve tasks that cannot be explicitly programmed, including pattern recognition, object detection, and autonomous control. To ensure the safe integration of ML in avionics, the European Union Aviation Safety Agency (EASA) proposed the W-shaped development process. However, it focuses on supervised learning, leaving aside Reinforcement Learning (RL), which is particularly suitable for decision-making problems. This paper presents a practical analysis of the applicability of the W-shaped process to systems that include RL components. To support this analysis, the W-shaped process is applied to a drone navigation and collision avoidance system under safety requirements on a real-time embedded platform. This paper demonstrates that processes related to learning must be adapted to support the interactive nature of RL, while those related to implementation are agnostic to the ML technique and thus compatible with RL. This paper also proposes concrete adaptations, including scenario-based data management, a reinterpretation of data quality requirements, and the explicit definition of the agent-environment interface before training.
Authors
- Ángel-Grover Pérez-Muñoz (ORCID: https://orcid.org/0009-0007-0760-6071)
- A. Alonso (ORCID: https://orcid.org/0000-0003-1259-0573)
- Hernán García-Quijano
- María S. Pérez (ORCID: https://orcid.org/0000-0003-2949-3307)
- Guillermo López-García
Institutions
- Universidad Politécnica de Madrid (ES)
Publication Details
- Journal
- Machine Learning with Applications
- Published
- 2026-09-18
- DOI
- https://doi.org/10.1016/j.mlwa.2026.101009
- Primary Topic
- Adversarial Robustness in Machine Learning
- Type
- article
- Field-Weighted Citation Impact
- 0.00