AutoRL: A Tightly Synchronized ROS2–Gazebo Pipeline for Offline-Trained Reinforcement Learning-Based Multirotor Attitude Control

This paper proposes AutoRL, a high-fidelity Robot Operating System 2 (ROS2)–Gazebo simulation pipeline that addresses a critical reproducibility gap in learning-based flight control: existing reinforcement learning (RL) frameworks for unmanned aerial vehicles (UAVs) lack deterministic, step-level coupling between control actions and physics updates. AutoRL enforces a strict one-to-one correspondence between agent actions and physics updates via blocking ROS2 service calls, preserving the Markov property required for stable policy learning and enabling verifiable reproducibility independent of the learning algorithm. A composite reward function jointly optimizes attitude tracking accuracy, oscillation suppression, actuator smoothness, and disturbance robustness. Its modular, service-oriented architecture provides a reusable framework for offline-trained RL research. A proximal policy optimization (PPO) controller trained within AutoRL validates the framework, demonstrating consistent convergence and stable performance across multiple independently seeded runs. Determinism was experimentally verified across two regimes: with Gaussian IMU noise disabled, repeated rollouts produced bit-identical trajectories, while with noise enabled, the measured distribution of trajectory divergence agreed with a reference distribution drawn from the declared sensor model, together confirming that the blocking service call architecture eliminates all non-stochastic sources of nondeterminism between the agent and the physics engine. The trained model is exported in a lightweight form compatible with embedded flight control firmware and remains adaptable across airframe configurations by automatically recomputing the control allocation matrix from configuration files. The total policy network contains only 10,628 trainable parameters, and inference was measured on a Cortex-M7 microcontroller at 598.7 µs per step, 15.0% of the 4 ms control period.

Authors

Institutions

Publication Details

Journal
Aerospace
Published
2026-09-10
DOI
https://doi.org/10.3390/aerospace13090825
Primary Topic
Aerospace and Aviation Technology
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

AutoRL: A Tightly Synchronized ROS2–Gazebo Pipeline for Offline-Trained Reinforcement Learning-Based Multirotor Attitude Control

Khaled Jarrah, Osamah Rawashdeh
Aerospace
Aerospace and Aviation Technology
article

AutoRL: A Tightly Synchronized ROS2–Gazebo Pipeline for Offline-Trained Reinforcement Learning-Based Multirotor Attitude Control

Khaled Jarrah, Osamah Rawashdeh
article en

Abstract

This paper proposes AutoRL, a high-fidelity Robot Operating System 2 (ROS2)–Gazebo simulation pipeline that addresses a critical reproducibility gap in learning-based flight control: existing reinforcement learning (RL) frameworks for unmanned aerial vehicles (UAVs) lack deterministic, step-level coupling between control actions and physics updates. AutoRL enforces a strict one-to-one correspondence between agent actions and physics updates via blocking ROS2 service calls, preserving the Markov property required for stable policy learning and enabling verifiable reproducibility independent of the learning algorithm. A composite reward function jointly optimizes attitude tracking accuracy, oscillation suppression, actuator smoothness, and disturbance robustness. Its modular, service-oriented architecture provides a reusable framework for offline-trained RL research. A proximal policy optimization (PPO) controller trained within AutoRL validates the framework, demonstrating consistent convergence and stable performance across multiple independently seeded runs. Determinism was experimentally verified across two regimes: with Gaussian IMU noise disabled, repeated rollouts produced bit-identical trajectories, while with noise enabled, the measured distribution of trajectory divergence agreed with a reference distribution drawn from the declared sensor model, together confirming that the blocking service call architecture eliminates all non-stochastic sources of nondeterminism between the agent and the physics engine. The trained model is exported in a lightweight form compatible with embedded flight control firmware and remains adaptable across airframe configurations by automatically recomputing the control allocation matrix from configuration files. The total policy network contains only 10,628 trainable parameters, and inference was measured on a Cortex-M7 microcontroller at 598.7 µs per step, 15.0% of the 4 ms control period.

AerospaceVol. 13(9)
Oakland University (US)
Openalex Percentile: Top 7%
Aerospace and Aviation Technology
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

AutoRL: A Tightly Synchronized ROS2–Gazebo Pipeline for Offline-Trained Reinforcement Learning-Based Multirotor Attitude Control — Khaled Jarrah, Osamah Rawashdeh · Aerospace (2026) | TGRS Research Map | TGRS