Model-Based Reinforcement Learning for HVAC Energy Optimization Under Hot, Mixed, and Cool Climates

HVAC control trades energy against thermal comfort, complicated by two building features: thermal mass spreads a setpoint change over hours, and the input-to-outcome mapping shifts across the year. Model-free algorithms such as PPO, SAC, and TD3 carry no model of building dynamics and cannot evaluate a setpoint’s downstream effect. We apply a Latent Dynamics Learning and Planning (LDLP) framework that learns a latent model of the building’s thermal response and plans with Monte Carlo Tree Search inside it. We also gate a reward formulation common in prior work, applying its comfort penalty only while the building is occupied. On Sinergym’s 5Zone environment under hot, mixed, and cool climates, LDLP is evaluated against PPO, SAC, TD3 and Sampled EfficientZero, an independent model-based controller run at the same budget. With energy normalized for the comfort achieved, LDLP consumes 4% to 26% less than PPO and SAC under the standard reward and 1% to 19% less under the gated reward. Running the same model with one simulation per decision, which removes planning, multiplies its normalized energy fivefold. Under the gated reward the deterministic TD3 policy degenerates onto a few fixed setpoints, so we report action diversity alongside the conventional metrics.

Authors

Institutions

Publication Details

Journal
Applied Sciences
Published
2026-09-15
DOI
https://doi.org/10.3390/app16189131
Primary Topic
Building Energy and Comfort Optimization
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Model-Based Reinforcement Learning for HVAC Energy Optimization Under Hot, Mixed, and Cool Climates

Chengnan Lu, Jinho Park
Applied Sciences
Building Energy and Comfort Optimization
article

Model-Based Reinforcement Learning for HVAC Energy Optimization Under Hot, Mixed, and Cool Climates

Chengnan Lu, Jinho Park
article en

Abstract

HVAC control trades energy against thermal comfort, complicated by two building features: thermal mass spreads a setpoint change over hours, and the input-to-outcome mapping shifts across the year. Model-free algorithms such as PPO, SAC, and TD3 carry no model of building dynamics and cannot evaluate a setpoint’s downstream effect. We apply a Latent Dynamics Learning and Planning (LDLP) framework that learns a latent model of the building’s thermal response and plans with Monte Carlo Tree Search inside it. We also gate a reward formulation common in prior work, applying its comfort penalty only while the building is occupied. On Sinergym’s 5Zone environment under hot, mixed, and cool climates, LDLP is evaluated against PPO, SAC, TD3 and Sampled EfficientZero, an independent model-based controller run at the same budget. With energy normalized for the comfort achieved, LDLP consumes 4% to 26% less than PPO and SAC under the standard reward and 1% to 19% less under the gated reward. Running the same model with one simulation per decision, which removes planning, multiplies its normalized energy fivefold. Under the gated reward the deterministic TD3 policy degenerates onto a few fixed setpoints, so we report action diversity alongside the conventional metrics.

Applied SciencesVol. 16(18)
Soongsil University (KR)
Climate action
Openalex Percentile: Top 14%
Building Energy and Comfort Optimization
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.