A single safe policy for fast charging control of multi-chemistry and multi-size battery packs using contextual reinforcement learning

Fast charging of lithium-ion battery (LIB) packs is constrained by thermal runaway risk, cell state-of-charge (SOC) imbalance, and accelerated capacity fade, challenges compounded by spatial temperature gradients and the diversity of cathode chemistries and configurations across electric vehicles and stationary storage. Existing reinforcement learning (RL) approaches address fast charging for a single chemistry and fixed pack size, requiring complete retraining when the battery type or configuration changes, which represents a critical barrier to scalable battery management system (BMS) deployment. This paper proposes a contextual Markov Decision Process (MDP) framework that encodes battery chemistry and pack size as explicit episode-level context variables. A single proximal policy optimisation (PPO) agent thereby generalises across three cathode chemistries, namely nickel manganese cobalt oxide (NMC), lithium cobalt oxide (LCO), and lithium iron phosphate (LFP), as well as five pack layouts (4–16 cells) without retraining. Seven neural surrogate models: three electro-thermal surrogates, three degradation models, and one thermal-runaway safety classifier, are trained on publicly available datasets and drive a configurable pack simulator, eliminating proprietary equivalent-circuit model identification. LFP’s near-flat open-circuit voltage plateau, rendering SOC unobservable from terminal voltage over 20–80% SOC, is resolved via Coulomb-counting state augmentation. A novel entropy-clipping mechanism prevents catastrophic PPO policy degradation in long-horizon episodes, and a two-phase progressive curriculum reduces LFP training cost by threefold. Evaluation across all 15 chemistry–layout combinations achieves 100% charging success (SOC: 0.20 → 0.80 , 75/75 episodes). Relative to 1 C CC–CV, the policy reduces charging time by 44% for NMC, 60% for LFP, and 79% for LCO, while peak temperatures remain below the corresponding chemistry-specific safety limits. Code and implementation details are available at GitHub repository .

Authors

Institutions

Publication Details

Journal
Applied Energy
Published
2026-10-06
DOI
https://doi.org/10.1016/j.apenergy.2026.128970
Primary Topic
Advanced Battery Technologies Research
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

A single safe policy for fast charging control of multi-chemistry and multi-size battery packs using contextual reinforcement learning

Pallavi Bharadwaj, Saumya Karan
Applied Energy
Advanced Battery Technologies Research
article

A single safe policy for fast charging control of multi-chemistry and multi-size battery packs using contextual reinforcement learning

Pallavi Bharadwaj, Saumya Karan
article en

Abstract

Fast charging of lithium-ion battery (LIB) packs is constrained by thermal runaway risk, cell state-of-charge (SOC) imbalance, and accelerated capacity fade, challenges compounded by spatial temperature gradients and the diversity of cathode chemistries and configurations across electric vehicles and stationary storage. Existing reinforcement learning (RL) approaches address fast charging for a single chemistry and fixed pack size, requiring complete retraining when the battery type or configuration changes, which represents a critical barrier to scalable battery management system (BMS) deployment. This paper proposes a contextual Markov Decision Process (MDP) framework that encodes battery chemistry and pack size as explicit episode-level context variables. A single proximal policy optimisation (PPO) agent thereby generalises across three cathode chemistries, namely nickel manganese cobalt oxide (NMC), lithium cobalt oxide (LCO), and lithium iron phosphate (LFP), as well as five pack layouts (4–16 cells) without retraining. Seven neural surrogate models: three electro-thermal surrogates, three degradation models, and one thermal-runaway safety classifier, are trained on publicly available datasets and drive a configurable pack simulator, eliminating proprietary equivalent-circuit model identification. LFP’s near-flat open-circuit voltage plateau, rendering SOC unobservable from terminal voltage over 20–80% SOC, is resolved via Coulomb-counting state augmentation. A novel entropy-clipping mechanism prevents catastrophic PPO policy degradation in long-horizon episodes, and a two-phase progressive curriculum reduces LFP training cost by threefold. Evaluation across all 15 chemistry–layout combinations achieves 100% charging success (SOC: 0.20 → 0.80 , 75/75 episodes). Relative to 1 C CC–CV, the policy reduces charging time by 44% for NMC, 60% for LFP, and 79% for LCO, while peak temperatures remain below the corresponding chemistry-specific safety limits. Code and implementation details are available at GitHub repository .

Applied EnergyVol. 427
Indian Institute of Technology Gandhinagar (IN)
Openalex Percentile: Top 21%
Advanced Battery Technologies Research
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.