AutoRL: A Tightly Synchronized ROS2–Gazebo Pipeline for Offline-Trained Reinforcement Learning-Based Multirotor Attitude Control
This paper proposes AutoRL, a high-fidelity Robot Operating System 2 (ROS2)–Gazebo simulation pipeline that addresses a critical reproducibility gap in learning-based flight control: existing reinforcement learning (RL) frameworks for unmanned aerial vehicles (UAVs) lack deterministic, step-level coupling between control actions and physics updates. AutoRL enforces a strict one-to-one correspondence between agent actions and physics updates via blocking ROS2 service calls, preserving the Markov property required for stable policy learning and enabling verifiable reproducibility independent of the learning algorithm. A composite reward function jointly optimizes attitude tracking accuracy, oscillation suppression, actuator smoothness, and disturbance robustness. Its modular, service-oriented architecture provides a reusable framework for offline-trained RL research. A proximal policy optimization (PPO) controller trained within AutoRL validates the framework, demonstrating consistent convergence and stable performance across multiple independently seeded runs. Determinism was experimentally verified across two regimes: with Gaussian IMU noise disabled, repeated rollouts produced bit-identical trajectories, while with noise enabled, the measured distribution of trajectory divergence agreed with a reference distribution drawn from the declared sensor model, together confirming that the blocking service call architecture eliminates all non-stochastic sources of nondeterminism between the agent and the physics engine. The trained model is exported in a lightweight form compatible with embedded flight control firmware and remains adaptable across airframe configurations by automatically recomputing the control allocation matrix from configuration files. The total policy network contains only 10,628 trainable parameters, and inference was measured on a Cortex-M7 microcontroller at 598.7 µs per step, 15.0% of the 4 ms control period.
Authors
- Khaled Jarrah (ORCID: https://orcid.org/0000-0002-0124-2082)
- Osamah Rawashdeh (ORCID: https://orcid.org/0000-0002-2902-6365)
Institutions
- Oakland University (US)
Publication Details
- Journal
- Aerospace
- Published
- 2026-09-10
- DOI
- https://doi.org/10.3390/aerospace13090825
- Primary Topic
- Aerospace and Aviation Technology
- Type
- article
- Field-Weighted Citation Impact
- 0.00