Adaptive Heuristic Reinforcement Learning (AH-RL) for Safety-Aware Control in Building Energy Management Applications
Reinforcement learning has shown strong potential for sequential control problems, providing an effective framework for adaptive policy derivation in complex and dynamic environments. However, standalone RL controllers may behave unstably or too aggressively, particularly in safety-critical applications where constraint violations or risky actions are not acceptable. This issue is especially relevant in energy and building control, where performance optimization must be balanced with safe and reliable operation. The current work proposes Adaptive Heuristic Reinforcement Learning (AH-RL), a hybrid control framework that combines a baseline heuristic controller with a residual reinforcement learning policy through an adaptive, safety-aware blending mechanism. In this structure, the heuristic controller serves as a stable operational reference, while the RL component learns corrective actions to improve performance without causing catastrophic deviations. The proposed AH-RL method is evaluated across three CityLearn benchmark scenarios using SAC, DDPG, and TD3 as continuous-control RL backbones. The AH-RL variants are compared with the concerned rule-based controllers, the corresponding standalone RL controllers, and also constrained Lagrangian RL controllers. The results show that AH-RL consistently provides a substantially improved thermal-comfort profile compared with both standalone RL and Lagrangian-constrained RL, markedly reducing hot-side discomfort and limiting the severe overheating observed in these learned baselines. Most importantly, these thermal comfort improvements were not achieved by compromising control performance—instead, they were accompanied by an increase in the overall reward: According to evaluation, the AH-RL algorithms consistently improved the mean total reward in comparison to rule-based control and generally outperformed both the corresponding standalone RL, as well as the more advantageous Lagrangian-constrained controllers. Overall, the AH-RL approach was capable of providing a significantly more balanced reward–thermal-comfort trade-off without relying on excessive thermal degradation to achieve performance gains. In addition, no critical state excessive-deviation or action-variation events were observed under the selected diagnostic thresholds. Such findings highlight the potential of AH-RL for utilization in safety-sensitive energy management applications and support its potential deployment in real-world settings.
Authors
- Federico Minelli (ORCID: https://orcid.org/0000-0002-5045-6474)
- Hasan Hüseyin Çoban (ORCID: https://orcid.org/0000-0002-5284-0568)
- Iakovos Michailidis (ORCID: https://orcid.org/0000-0001-7295-8806)
- Elias B. Kosmatopoulos (ORCID: https://orcid.org/0000-0002-3735-4238)
- Panagiotis Michailidis (ORCID: https://orcid.org/0000-0001-7148-7010)
Institutions
- Democritus University of Thrace (GR)
- Bartin University (TR)
- Centre for Research and Technology Hellas (GR)
- University of Naples Federico II (IT)
Publication Details
- Journal
- Eng—Advances in Engineering
- Published
- 2026-09-01
- DOI
- https://doi.org/10.3390/eng7090438
- Primary Topic
- Building Energy and Comfort Optimization
- Type
- article
- Field-Weighted Citation Impact
- 0.00
Funders
- European Commission