Transformer-Based Temporal Reinforcement Learning for Series Elastic Actuator Control
Series Elastic Actuators (SEAs) have been widely adopted in collaborative robots, rehabilitation robots, and other compliant robotic systems owing to their excellent compliance and force-control performance. However, the introduction of elastic elements also increases dynamic complexity, giving rise to nonlinearities, parameter uncertainties, and oscillations. To address these issues, this paper proposes a Transformer-based temporal reinforcement learning control method for SEAs within the Proximal Policy Optimization (PPO) framework. First, a Transformer replaces the conventional Multi-Layer Perceptron (MLP) policy network to extract temporal dependency features from observation sequences via the self-attention mechanism. Second, a fixed-length historical state window is incorporated into the policy input to mitigate the performance degradation caused by partial observability (POMDP). Both simulation and hardware experiments validate the proposed method. In simulation, it achieves superior transient performance (a 10–90% rise time of 0.111 s and an overshoot of 3.68%) together with stable sinusoidal tracking. Real-world tests on a physical SEA platform further confirm its practical feasibility: multi-target step tracking attains a 10–90% rise time of approximately 0.15 s with overshoot below 2.3%, whereas long-duration sinusoidal tracking (60° amplitude, 0.1 Hz) maintains an RMS error of approximately 2.3° and smooth, chatter-free control actions. Comparative studies against MLP, LSTM, and CNN architectures reveal that the Transformer policy offers significant advantages in suppressing elastic oscillations and generating smooth control commands, confirming the efficacy of temporal modeling for the complex dynamics of SEA systems. An ablation study over the historical window length (L = 1–8) shows that at least four history steps are required to suppress elastic oscillations under fast target switching, while the mean inference time remains below 0.81 ms for L ≤ 8. In hardware payload experiments (0–1200 g), the tracking RMSE increases monotonically from 7.70° to 8.55° and the peak overshoot from 2.50% to 8.53%, while all runs remain stable without controller fault or divergence; the complete control task, including Transformer inference, executes within 0.61 ms on the real-time controller.
Authors
- Feiyan Min (ORCID: https://orcid.org/0000-0002-8610-2643)
- Xinyu Liu (ORCID: https://orcid.org/0009-0004-6776-1821)
- Ziqian Li (ORCID: https://orcid.org/0000-0002-9228-9643)
- Yaoyao Lu
- Yuan Liu
- Xiangzhong Yan
- Wenxuan Wu
Institutions
- Jinan University (CN)
- Wuhan Ship Development & Design Institute (CN)
- Beijing Computing Center (CN)
Publication Details
- Journal
- Actuators
- Published
- 2026-09-24
- DOI
- https://doi.org/10.3390/act15100504
- Primary Topic
- Prosthetics and Rehabilitation Robotics
- Type
- article
- Field-Weighted Citation Impact
- 0.00