Transformer-Based Temporal Reinforcement Learning for Series Elastic Actuator Control

Series Elastic Actuators (SEAs) have been widely adopted in collaborative robots, rehabilitation robots, and other compliant robotic systems owing to their excellent compliance and force-control performance. However, the introduction of elastic elements also increases dynamic complexity, giving rise to nonlinearities, parameter uncertainties, and oscillations. To address these issues, this paper proposes a Transformer-based temporal reinforcement learning control method for SEAs within the Proximal Policy Optimization (PPO) framework. First, a Transformer replaces the conventional Multi-Layer Perceptron (MLP) policy network to extract temporal dependency features from observation sequences via the self-attention mechanism. Second, a fixed-length historical state window is incorporated into the policy input to mitigate the performance degradation caused by partial observability (POMDP). Both simulation and hardware experiments validate the proposed method. In simulation, it achieves superior transient performance (a 10–90% rise time of 0.111 s and an overshoot of 3.68%) together with stable sinusoidal tracking. Real-world tests on a physical SEA platform further confirm its practical feasibility: multi-target step tracking attains a 10–90% rise time of approximately 0.15 s with overshoot below 2.3%, whereas long-duration sinusoidal tracking (60° amplitude, 0.1 Hz) maintains an RMS error of approximately 2.3° and smooth, chatter-free control actions. Comparative studies against MLP, LSTM, and CNN architectures reveal that the Transformer policy offers significant advantages in suppressing elastic oscillations and generating smooth control commands, confirming the efficacy of temporal modeling for the complex dynamics of SEA systems. An ablation study over the historical window length (L = 1–8) shows that at least four history steps are required to suppress elastic oscillations under fast target switching, while the mean inference time remains below 0.81 ms for L ≤ 8. In hardware payload experiments (0–1200 g), the tracking RMSE increases monotonically from 7.70° to 8.55° and the peak overshoot from 2.50% to 8.53%, while all runs remain stable without controller fault or divergence; the complete control task, including Transformer inference, executes within 0.61 ms on the real-time controller.

Authors

Institutions

Publication Details

Journal
Actuators
Published
2026-09-24
DOI
https://doi.org/10.3390/act15100504
Primary Topic
Prosthetics and Rehabilitation Robotics
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Transformer-Based Temporal Reinforcement Learning for Series Elastic Actuator Control

Feiyan Min, Xinyu Liu, Ziqian Li, Yaoyao Lu et al.
Actuators
Prosthetics and Rehabilitation Robotics
article

Transformer-Based Temporal Reinforcement Learning for Series Elastic Actuator Control

Feiyan Min, Xinyu Liu, Ziqian Li, Yaoyao Lu, Yuan Liu, Xiangzhong Yan, Wenxuan Wu
article en

Abstract

Series Elastic Actuators (SEAs) have been widely adopted in collaborative robots, rehabilitation robots, and other compliant robotic systems owing to their excellent compliance and force-control performance. However, the introduction of elastic elements also increases dynamic complexity, giving rise to nonlinearities, parameter uncertainties, and oscillations. To address these issues, this paper proposes a Transformer-based temporal reinforcement learning control method for SEAs within the Proximal Policy Optimization (PPO) framework. First, a Transformer replaces the conventional Multi-Layer Perceptron (MLP) policy network to extract temporal dependency features from observation sequences via the self-attention mechanism. Second, a fixed-length historical state window is incorporated into the policy input to mitigate the performance degradation caused by partial observability (POMDP). Both simulation and hardware experiments validate the proposed method. In simulation, it achieves superior transient performance (a 10–90% rise time of 0.111 s and an overshoot of 3.68%) together with stable sinusoidal tracking. Real-world tests on a physical SEA platform further confirm its practical feasibility: multi-target step tracking attains a 10–90% rise time of approximately 0.15 s with overshoot below 2.3%, whereas long-duration sinusoidal tracking (60° amplitude, 0.1 Hz) maintains an RMS error of approximately 2.3° and smooth, chatter-free control actions. Comparative studies against MLP, LSTM, and CNN architectures reveal that the Transformer policy offers significant advantages in suppressing elastic oscillations and generating smooth control commands, confirming the efficacy of temporal modeling for the complex dynamics of SEA systems. An ablation study over the historical window length (L = 1–8) shows that at least four history steps are required to suppress elastic oscillations under fast target switching, while the mean inference time remains below 0.81 ms for L ≤ 8. In hardware payload experiments (0–1200 g), the tracking RMSE increases monotonically from 7.70° to 8.55° and the peak overshoot from 2.50% to 8.53%, while all runs remain stable without controller fault or divergence; the complete control task, including Transformer inference, executes within 0.61 ms on the real-time controller.

ActuatorsVol. 15(10)
Jinan University (CN), Wuhan Ship Development & Design Institute (CN), Beijing Computing Center (CN)
Life below water
Openalex Percentile: Top 21%
Prosthetics and Rehabilitation Robotics
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.