Adaptive control of twin rotor system based on two-stage reinforcement learning with quantile-guided trust region

Operating-condition changes can degrade dual-channel active disturbance rejection control (ADRC) on a twin-rotor system. This study implements an artificial intelligence method, Quantile-Guided Trust Region Adaptive Control through Two-Stage Soft Actor-Critic-to-Deep Deterministic Policy Gradient Learning (Q-TRACT), for online ADRC self-tuning. Return-gated soft actor-critic (SAC) actions bound a 12 -dimensional parameter domain through componentwise 0.1 − 0.9 empirical quantiles; projected deep deterministic policy gradient (DDPG) learns within it. Union Projection enforces an undershoot-informed feasible voltage set before actuation. Archive-admission diagnostics, order-statistic sensitivity, and a conditional local model-to-parameter mismatch bound characterize transfer. A coupled-model audit verifies a strict common-Lyapunov certificate for a post hoc screened core. All five independent runs achieved two-channel settling; control-metric coefficients of variation were 7 % − 12 % . In a fixed-archive, fixed-seed comparison, untrimmed bounds increase pitch/yaw settling times by 59.2 % / 60.0 % and integral absolute error (IAE) by 48.7 % / 53.2 % relative to 0.1 − 0.9 bounds. After single-condition training, the selected policy was frozen and evaluated on a 100 -combination target grid consisting of 99 unseen joint pairs plus the training pair, and on complex references, disturbances, and Quanser Aero experiments. Relative to an architecture-matched fixed-parameter ADRC, Q-TRACT reduces the largest pitch and yaw steady-state errors by 96.8 % and 97.5 % , respectively, and the largest yaw overshoot by 87.1 % , although its largest pitch overshoot is 44.7 % higher. Ramp experiments show 49.5 % − 77.6 % lower IAE than Manual Tuning. Crossed ablations show no universal metric dominance and a controller-dependent effect of Union Projection. Q-TRACT provides ADRC self-tuning with execution-layer input admissibility across the tested conditions.

Authors

Institutions

Publication Details

Journal
Engineering Applications of Artificial Intelligence
Published
2026-09-24
DOI
https://doi.org/10.1016/j.engappai.2026.116307
Primary Topic
Plasma and Flow Control in Aerodynamics
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Adaptive control of twin rotor system based on two-stage reinforcement learning with quantile-guided trust region

Wei Wei, Donghai Li, Min Zhu, Song Huang
Engineering Applications of Artificial Intelligence
Plasma and Flow Control in Aerodynamics
article

Adaptive control of twin rotor system based on two-stage reinforcement learning with quantile-guided trust region

Wei Wei, Donghai Li, Min Zhu, Song Huang
article en

Abstract

Operating-condition changes can degrade dual-channel active disturbance rejection control (ADRC) on a twin-rotor system. This study implements an artificial intelligence method, Quantile-Guided Trust Region Adaptive Control through Two-Stage Soft Actor-Critic-to-Deep Deterministic Policy Gradient Learning (Q-TRACT), for online ADRC self-tuning. Return-gated soft actor-critic (SAC) actions bound a 12 -dimensional parameter domain through componentwise 0.1 − 0.9 empirical quantiles; projected deep deterministic policy gradient (DDPG) learns within it. Union Projection enforces an undershoot-informed feasible voltage set before actuation. Archive-admission diagnostics, order-statistic sensitivity, and a conditional local model-to-parameter mismatch bound characterize transfer. A coupled-model audit verifies a strict common-Lyapunov certificate for a post hoc screened core. All five independent runs achieved two-channel settling; control-metric coefficients of variation were 7 % − 12 % . In a fixed-archive, fixed-seed comparison, untrimmed bounds increase pitch/yaw settling times by 59.2 % / 60.0 % and integral absolute error (IAE) by 48.7 % / 53.2 % relative to 0.1 − 0.9 bounds. After single-condition training, the selected policy was frozen and evaluated on a 100 -combination target grid consisting of 99 unseen joint pairs plus the training pair, and on complex references, disturbances, and Quanser Aero experiments. Relative to an architecture-matched fixed-parameter ADRC, Q-TRACT reduces the largest pitch and yaw steady-state errors by 96.8 % and 97.5 % , respectively, and the largest yaw overshoot by 87.1 % , although its largest pitch overshoot is 44.7 % higher. Ramp experiments show 49.5 % − 77.6 % lower IAE than Manual Tuning. Crossed ablations show no universal metric dominance and a controller-dependent effect of Union Projection. Q-TRACT provides ADRC self-tuning with execution-layer input admissibility across the tested conditions.

Engineering Applications of Artificial IntelligenceVol. 184
Beijing University of Posts and Telecommunications (CN), Tsinghua University (CN)
Peace, Justice and strong institutions
Openalex Percentile: Top 8%
Plasma and Flow Control in Aerodynamics
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.