Model-Based Reinforcement Learning for HVAC Energy Optimization Under Hot, Mixed, and Cool Climates
HVAC control trades energy against thermal comfort, complicated by two building features: thermal mass spreads a setpoint change over hours, and the input-to-outcome mapping shifts across the year. Model-free algorithms such as PPO, SAC, and TD3 carry no model of building dynamics and cannot evaluate a setpoint’s downstream effect. We apply a Latent Dynamics Learning and Planning (LDLP) framework that learns a latent model of the building’s thermal response and plans with Monte Carlo Tree Search inside it. We also gate a reward formulation common in prior work, applying its comfort penalty only while the building is occupied. On Sinergym’s 5Zone environment under hot, mixed, and cool climates, LDLP is evaluated against PPO, SAC, TD3 and Sampled EfficientZero, an independent model-based controller run at the same budget. With energy normalized for the comfort achieved, LDLP consumes 4% to 26% less than PPO and SAC under the standard reward and 1% to 19% less under the gated reward. Running the same model with one simulation per decision, which removes planning, multiplies its normalized energy fivefold. Under the gated reward the deterministic TD3 policy degenerates onto a few fixed setpoints, so we report action diversity alongside the conventional metrics.
Authors
- Chengnan Lu
- Jinho Park (ORCID: https://orcid.org/0000-0002-8694-2976)
Institutions
- Soongsil University (KR)
Publication Details
- Journal
- Applied Sciences
- Published
- 2026-09-15
- DOI
- https://doi.org/10.3390/app16189131
- Primary Topic
- Building Energy and Comfort Optimization
- Type
- article
- Field-Weighted Citation Impact
- 0.00