Constraint-Aware and Energy-Efficient Control of an Integrated Chemical Process via Offline-to-Online Reinforcement Learning
Abstract Integrated chemical processes involving reaction, separation, recycle, and heat integration often exhibit strong nonlinearities, unit-to-unit coupling, and multiple operating constraints, making their safe and efficient control challenging. Purely online reinforcement learning may require risky trial-and-error exploration, whereas purely offline learning can suffer from distribution shift and limited adaptability. To address these issues, an offline-to-online reinforcement learning framework is developed for an integrated chemical process with two nonisothermal reaction stages, membrane-assisted dehydration, separation, recycle, and heat recovery. In the offline stage, Adaptive Safe Conservative Q-Learning (AS-CQL) combines state-dependent conservatism with uncertainty-aware safety value estimation to obtain a reliable initial policy from historical data. The pretrained policy is then transferred to Twin Delayed Deep Deterministic Policy Gradient (TD3) for online adaptation. The control objective considers target-product tracking, process constraints, and smooth control manipulation. Simulations show that AS-CQL + TD3 achieves favorable tracking and constraint-handling performance and the lowest cumulative cost among the compared controllers, demonstrating its effectiveness for data-driven control of integrated chemical processes.
Authors
- Jun Rao (ORCID: https://orcid.org/0000-0002-7467-4898)
- Jingcheng Wang (ORCID: https://orcid.org/0000-0002-4277-1263)
- Chengtian Cui
- Daye Yang
Institutions
- Åbo Akademi University (FI)
- Shanghai Jiao Tong University (CN)
Publication Details
- Journal
- Industrial & Engineering Chemistry Research
- Published
- 2026-09-25
- DOI
- https://doi.org/10.1021/acs.iecr.6c02782
- Primary Topic
- Advanced Control Systems Optimization
- Type
- article
- Field-Weighted Citation Impact
- 0.00