Action-conditioned dynamics learning for model predictive control of flexible fish locomotion
Closed-loop navigation of an autonomous flexible fish requires rapid prediction of the trajectory evolution induced by body-wave commands. However, evaluating candidate command sequences one by one with high-fidelity flow simulations would incur prohibitive computational cost. To address this challenge, this work proposes an action-conditioned fish–fluid dynamics learning framework for predictive planning of flexible-fish locomotion. The framework uses disagreement-driven active sampling to selectively enrich the trajectory database in regions where the latent dynamics remain highly uncertain. Compared with random-action sampling, the guided database increases the mean and median normalized one-step response amplitudes by 19.6% and 22.5%, respectively, reduces the weak-response ratio by 22.7%, and enlarges the multivariate response covariance volume by approximately 1.76 times, thereby improving response-level diversity. Building on this enriched database, the framework further develops a rollout robust recurrent state space model (RR-RSSM) to suppress recursive prediction drift over the planning horizon. At a rollout horizon of k = 20 , RR-RSSM achieves a position RMSE of 14.327 grid units and a heading MAE of 3.286°, maintaining lower prediction errors than ARX, LSTM, and the Baseline RSSM. The learned dynamics model is then used to rapidly roll out the trajectory consequences of candidate body-wave command sequences and to plan swimming actions online, enabling closed-loop autonomous navigation of the fish-like swimmer. Overall, the results demonstrate that action-conditioned fish–fluid dynamics learning provides an efficient model-based planning approach for autonomous underwater locomotion. • Data-driven reduced-order modeling and MPC are developed for fish–fluid coupling. • Disagreement-guided sampling increases the response volume by 1.76-fold. • At K = 20 , RR-RSSM cuts position RMSE by 59.2% and heading MAE by 68.8%. • Learned-model MPC reaches different targets without retraining.
Authors
- Fengchen LI (ORCID: https://orcid.org/0000-0003-3271-7949)
- Chunyu Wang (ORCID: https://orcid.org/0000-0001-5165-7959)
- Xiaobin Li (ORCID: https://orcid.org/0000-0001-9297-6171)
- Hong-Na Zhang (ORCID: https://orcid.org/0000-0002-1161-0897)
- Yan Wang
Institutions
- Tianjin University (CN)
Publication Details
- Journal
- Ocean Engineering
- Published
- 2026-09-24
- DOI
- https://doi.org/10.1016/j.oceaneng.2026.128194
- Primary Topic
- Biomimetic flight and propulsion mechanisms
- Type
- article
- Field-Weighted Citation Impact
- 0.00