Mitigating Physically Implausible Energy Performance Forecasting in Residential Buildings Using Logarithmic Transformation-Based Machine Learning Analysis: A Case Study Using RECS and ResStock Datasets

Data-driven analysis of residential building energy consumption is essential for understanding how homes perform, improving model accuracy, and guiding large-scale energy efficiency decisions. In this study, we apply several machine learning methods, including CatBoost, LightGBM, Random Forest, XGBoost, and a Neural Network, to predict total energy consumption as well as space heating and cooling loads across two diverse datasets – the Energy Information Administration’s Residential Energy Consumption Survey (RECS) of nearly 18,500 real households, and the Department of Energy’s Residential Stock (ResStock) database, a compendium of nearly 550,000 statistically-developed building simulations. To address the significant skewness often observed in energy consumption distributions, we applied a logarithmic transformation to all energy performance metrics, including total energy consumption and space heating and cooling energy use. In its raw form, residential energy performance (whether total usage, space heating, or space cooling) often exhibits a strongly right-skewed distribution: a small fraction of homes accounts for disproportionately large energy demands. Such heavy-tailed data can bias statistical models, destabilize variance estimates, and undermine predictive accuracy. By applying a natural logarithm transformation (with an additive constant, e, to handle low or zero values), we compress extremely high values and stretch values near zero, yielding a distribution that is more symmetric and statistically tractable. After fitting our models in the transformed space, we back-transform predictions to the original energy scale for interpretation. This log-transformation approach enhances the statistical robustness of our models while preserving their physical interpretability, effectively mitigating physically implausible negative forecasts and improving predictive consistency across highly skewed energy variables. These data-driven models serve as fast, scalable predictors of residential energy performance for individual homes. By basing predictions on actual (RECS) and/or modeled consumption (ResStock), they improve the accuracy of national energy demand estimates and enable identification of potential energy efficiency upgrades. This approach enhances traditional tools with greater empirical predictive accuracy, particularly for skewed end uses such as cooling. Using the logarithmic transformation approach resulted in significant model performance improvement for cooling energy consumption forecasts, with R2 values rising from 0.867 to 0.928 and decreases in MAE and RMSE values from 5943 to 4517 and 9441 to 6911, respectively, thereby underscoring the value of this approach in eliminating physically implausible values in energy performance forecasting.

Authors

Publication Details

Journal
Digital Repository at the University of Maryland (University of Maryland College Park)
Published
2026-09-10
DOI
https://doi.org/10.13016/g4af-o722
Primary Topic
Building Energy and Comfort Optimization
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Mitigating Physically Implausible Energy Performance Forecasting in Residential Buildings Using Logarithmic Transformation-Based Machine Learning Analysis: A Case Study Using RECS and ResStock Datasets

Aditya Ramnarayan, Patricia Gunderson, Fatih Evren
Digital Repository at the University of Maryland (University of Maryland College Park)
Building Energy and Comfort Optimization
article

Mitigating Physically Implausible Energy Performance Forecasting in Residential Buildings Using Logarithmic Transformation-Based Machine Learning Analysis: A Case Study Using RECS and ResStock Datasets

Aditya Ramnarayan, Patricia Gunderson, Fatih Evren
article en

Abstract

Data-driven analysis of residential building energy consumption is essential for understanding how homes perform, improving model accuracy, and guiding large-scale energy efficiency decisions. In this study, we apply several machine learning methods, including CatBoost, LightGBM, Random Forest, XGBoost, and a Neural Network, to predict total energy consumption as well as space heating and cooling loads across two diverse datasets – the Energy Information Administration’s Residential Energy Consumption Survey (RECS) of nearly 18,500 real households, and the Department of Energy’s Residential Stock (ResStock) database, a compendium of nearly 550,000 statistically-developed building simulations. To address the significant skewness often observed in energy consumption distributions, we applied a logarithmic transformation to all energy performance metrics, including total energy consumption and space heating and cooling energy use. In its raw form, residential energy performance (whether total usage, space heating, or space cooling) often exhibits a strongly right-skewed distribution: a small fraction of homes accounts for disproportionately large energy demands. Such heavy-tailed data can bias statistical models, destabilize variance estimates, and undermine predictive accuracy. By applying a natural logarithm transformation (with an additive constant, e, to handle low or zero values), we compress extremely high values and stretch values near zero, yielding a distribution that is more symmetric and statistically tractable. After fitting our models in the transformed space, we back-transform predictions to the original energy scale for interpretation. This log-transformation approach enhances the statistical robustness of our models while preserving their physical interpretability, effectively mitigating physically implausible negative forecasts and improving predictive consistency across highly skewed energy variables. These data-driven models serve as fast, scalable predictors of residential energy performance for individual homes. By basing predictions on actual (RECS) and/or modeled consumption (ResStock), they improve the accuracy of national energy demand estimates and enable identification of potential energy efficiency upgrades. This approach enhances traditional tools with greater empirical predictive accuracy, particularly for skewed end uses such as cooling. Using the logarithmic transformation approach resulted in significant model performance improvement for cooling energy consumption forecasts, with R2 values rising from 0.867 to 0.928 and decreases in MAE and RMSE values from 5943 to 4517 and 9441 to 6911, respectively, thereby underscoring the value of this approach in eliminating physically implausible values in energy performance forecasting.

Digital Repository at the University of Maryland (University of Maryland College Park)
Affordable and clean energy
Openalex Percentile: Top 14%
Building Energy and Comfort Optimization
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.