A hierarchical deep reinforcement learning-model predictive control framework for lithium-ion battery fast charging under expansion-force constraints

Fast charging of lithium-ion batteries is constrained by the trade-off between charging efficiency and physical safety, especially when high-current operation narrows the admissible electrical, thermal, and mechanical margins. Traditional model-based controllers can explicitly handle constraints, but their performance depends on model fidelity and may become conservative near safety boundaries. In contrast, model-free deep reinforcement learning (DRL) provides flexible sequential decision-making but does not by itself guarantee hard safety-constraint satisfaction. This study presents a multi-stage constant-current (MCC)-inspired hierarchical DRL–model predictive control (MPC) framework that couples staged-current policy search with MPC-based safety correction under expansion-force constraints. Expansion force is incorporated as a dynamic safety constraint together with terminal voltage and surface temperature, allowing their coupled responses to be considered during control. A long short-term memory (LSTM)-based surrogate model is developed to predict battery-state evolution and support control-oriented optimization. Independent trajectory- and C-rate-holdout tests are used to examine its multi-step prediction error. Based on the surrogate predictions, the upper DRL layer selects candidate staged-current actions, whereas the lower MPC layer performs quadratic-programming (QP)-based safety correction before execution. Comparative simulation results against 3C constant-current (CC), native MPC, and pure DRL baselines show that the proposed method reaches 90% state of charge (SOC) in 728 s and improves the comprehensive score by 23.02%, 6.53%, and 11.64%, respectively. The reported trajectories remain below the prescribed voltage and temperature limits and the MPC expansion-force operating envelope.

Authors

Institutions

Publication Details

Journal
Journal of Energy Storage
Published
2026-10-09
DOI
https://doi.org/10.1016/j.est.2026.124691
Primary Topic
Advanced Battery Technologies Research
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

A hierarchical deep reinforcement learning-model predictive control framework for lithium-ion battery fast charging under expansion-force constraints

Lei Zhang, Kuijie Li, Shun Tang, Tianyou Wang
Journal of Energy Storage
Advanced Battery Technologies Research
article

A hierarchical deep reinforcement learning-model predictive control framework for lithium-ion battery fast charging under expansion-force constraints

Lei Zhang, Kuijie Li, Shun Tang, Tianyou Wang
article en

Abstract

Fast charging of lithium-ion batteries is constrained by the trade-off between charging efficiency and physical safety, especially when high-current operation narrows the admissible electrical, thermal, and mechanical margins. Traditional model-based controllers can explicitly handle constraints, but their performance depends on model fidelity and may become conservative near safety boundaries. In contrast, model-free deep reinforcement learning (DRL) provides flexible sequential decision-making but does not by itself guarantee hard safety-constraint satisfaction. This study presents a multi-stage constant-current (MCC)-inspired hierarchical DRL–model predictive control (MPC) framework that couples staged-current policy search with MPC-based safety correction under expansion-force constraints. Expansion force is incorporated as a dynamic safety constraint together with terminal voltage and surface temperature, allowing their coupled responses to be considered during control. A long short-term memory (LSTM)-based surrogate model is developed to predict battery-state evolution and support control-oriented optimization. Independent trajectory- and C-rate-holdout tests are used to examine its multi-step prediction error. Based on the surrogate predictions, the upper DRL layer selects candidate staged-current actions, whereas the lower MPC layer performs quadratic-programming (QP)-based safety correction before execution. Comparative simulation results against 3C constant-current (CC), native MPC, and pure DRL baselines show that the proposed method reaches 90% state of charge (SOC) in 728 s and improves the comprehensive score by 23.02%, 6.53%, and 11.64%, respectively. The reported trajectories remain below the prescribed voltage and temperature limits and the MPC expansion-force operating envelope.

Journal of Energy StorageVol. 182
Wuhan University of Science and Technology (CN), Huazhong University of Science and Technology (CN), Fuzhou University (CN)
Openalex Percentile: Top 21%
Advanced Battery Technologies Research
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.