Adaptive Heuristic Reinforcement Learning (AH-RL) for Safety-Aware Control in Building Energy Management Applications

Reinforcement learning has shown strong potential for sequential control problems, providing an effective framework for adaptive policy derivation in complex and dynamic environments. However, standalone RL controllers may behave unstably or too aggressively, particularly in safety-critical applications where constraint violations or risky actions are not acceptable. This issue is especially relevant in energy and building control, where performance optimization must be balanced with safe and reliable operation. The current work proposes Adaptive Heuristic Reinforcement Learning (AH-RL), a hybrid control framework that combines a baseline heuristic controller with a residual reinforcement learning policy through an adaptive, safety-aware blending mechanism. In this structure, the heuristic controller serves as a stable operational reference, while the RL component learns corrective actions to improve performance without causing catastrophic deviations. The proposed AH-RL method is evaluated across three CityLearn benchmark scenarios using SAC, DDPG, and TD3 as continuous-control RL backbones. The AH-RL variants are compared with the concerned rule-based controllers, the corresponding standalone RL controllers, and also constrained Lagrangian RL controllers. The results show that AH-RL consistently provides a substantially improved thermal-comfort profile compared with both standalone RL and Lagrangian-constrained RL, markedly reducing hot-side discomfort and limiting the severe overheating observed in these learned baselines. Most importantly, these thermal comfort improvements were not achieved by compromising control performance—instead, they were accompanied by an increase in the overall reward: According to evaluation, the AH-RL algorithms consistently improved the mean total reward in comparison to rule-based control and generally outperformed both the corresponding standalone RL, as well as the more advantageous Lagrangian-constrained controllers. Overall, the AH-RL approach was capable of providing a significantly more balanced reward–thermal-comfort trade-off without relying on excessive thermal degradation to achieve performance gains. In addition, no critical state excessive-deviation or action-variation events were observed under the selected diagnostic thresholds. Such findings highlight the potential of AH-RL for utilization in safety-sensitive energy management applications and support its potential deployment in real-world settings.

Authors

Institutions

Publication Details

Journal
Eng—Advances in Engineering
Published
2026-09-01
DOI
https://doi.org/10.3390/eng7090438
Primary Topic
Building Energy and Comfort Optimization
Type
article
Field-Weighted Citation Impact
0.00

Funders

Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Adaptive Heuristic Reinforcement Learning (AH-RL) for Safety-Aware Control in Building Energy Management Applications

Federico Minelli, Hasan Hüseyin Çoban, Iakovos Michailidis, Elias B. Kosmatopoulos et al.
Eng—Advances in Engineering
Building Energy and Comfort Optimization
article

Adaptive Heuristic Reinforcement Learning (AH-RL) for Safety-Aware Control in Building Energy Management Applications

Federico Minelli, Hasan Hüseyin Çoban, Iakovos Michailidis, Elias B. Kosmatopoulos, Panagiotis Michailidis
article en

Abstract

Reinforcement learning has shown strong potential for sequential control problems, providing an effective framework for adaptive policy derivation in complex and dynamic environments. However, standalone RL controllers may behave unstably or too aggressively, particularly in safety-critical applications where constraint violations or risky actions are not acceptable. This issue is especially relevant in energy and building control, where performance optimization must be balanced with safe and reliable operation. The current work proposes Adaptive Heuristic Reinforcement Learning (AH-RL), a hybrid control framework that combines a baseline heuristic controller with a residual reinforcement learning policy through an adaptive, safety-aware blending mechanism. In this structure, the heuristic controller serves as a stable operational reference, while the RL component learns corrective actions to improve performance without causing catastrophic deviations. The proposed AH-RL method is evaluated across three CityLearn benchmark scenarios using SAC, DDPG, and TD3 as continuous-control RL backbones. The AH-RL variants are compared with the concerned rule-based controllers, the corresponding standalone RL controllers, and also constrained Lagrangian RL controllers. The results show that AH-RL consistently provides a substantially improved thermal-comfort profile compared with both standalone RL and Lagrangian-constrained RL, markedly reducing hot-side discomfort and limiting the severe overheating observed in these learned baselines. Most importantly, these thermal comfort improvements were not achieved by compromising control performance—instead, they were accompanied by an increase in the overall reward: According to evaluation, the AH-RL algorithms consistently improved the mean total reward in comparison to rule-based control and generally outperformed both the corresponding standalone RL, as well as the more advantageous Lagrangian-constrained controllers. Overall, the AH-RL approach was capable of providing a significantly more balanced reward–thermal-comfort trade-off without relying on excessive thermal degradation to achieve performance gains. In addition, no critical state excessive-deviation or action-variation events were observed under the selected diagnostic thresholds. Such findings highlight the potential of AH-RL for utilization in safety-sensitive energy management applications and support its potential deployment in real-world settings.

Eng—Advances in EngineeringVol. 7(9)
Democritus University of Thrace (GR), Bartin University (TR), Centre for Research and Technology Hellas (GR), University of Naples Federico II (IT)
European Commission
Affordable and clean energy
Openalex Percentile: Top 14%
Building Energy and Comfort Optimization
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.