Benchmarking Model Complexity for Short-Term Surrogate Forecasting of High-Resolution Urban WRF Outputs During an Extreme Heatwave in Chongqing

High-resolution urban Weather Research and Forecasting (WRF) simulations resolve interactions among urban form, near-surface meteorology, atmospheric dynamics, and building-energy processes, but their computational cost limits repeated scenario analysis. This study compares eight pointwise temporal surrogates (Persistence, previous-day same-hour, Ridge, Random Forest, XGBoost, multilayer perceptron (MLP), gated recurrent unit (GRU), and long short-term memory (LSTM)) for recursive 24 h prediction of T2, Q2, W10, and WRF-BEM diagnostic air-conditioning power density (AC) from a 120 h WRF-BEP/BEM-LCZ heatwave simulation over Chongqing. Model development used three expanding-window folds, fold-specific scaling, fixed monitoring cells, and five fixed seeds for stochastic families; the h97–h120 full-domain test remained untouched until all modeling choices had been finalized. Station comparisons provided a limited check on the physical plausibility of the WRF meteorological fields; AC was not validated against metered energy use. Model performance depended strongly on the target variable. Random Forest performed best for T2 (1.081 ± 0.002 °C), LSTM for Q2 (0.969 ± 0.063 g/kg), and MLP for AC (0.505 ± 0.187 W/m2), whereas deterministic Ridge was best for W10 (1.095 m/s). W10 had weaker temporal memory (median lag-1/lag-24 ACF 0.581/0.211). Ridge outperformed the Previous-day baseline for 75.0% of grid cells, although errors increased sharply in the top wind-speed decile and during rapid changes. For this single-event, single-city benchmark, recurrent models did not improve W10 accuracy over Ridge. Model choice depended on the target and the balance between accuracy, spatial fidelity, and computational cost; transfer to other events and meteorological regimes remains untested.

Authors

Institutions

Publication Details

Journal
Land
Published
2026-09-16
DOI
https://doi.org/10.3390/land15091726
Primary Topic
Urban Heat Island Mitigation
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Benchmarking Model Complexity for Short-Term Surrogate Forecasting of High-Resolution Urban WRF Outputs During an Extreme Heatwave in Chongqing

Bao‐Jie He, Yanan Liu, Hong Li, Xie Runjie et al.
Land
Urban Heat Island Mitigation
article

Benchmarking Model Complexity for Short-Term Surrogate Forecasting of High-Resolution Urban WRF Outputs During an Extreme Heatwave in Chongqing

Bao‐Jie He, Yanan Liu, Hong Li, Xie Runjie, Maoyuan Chai, Ruiqing Du
article en

Abstract

High-resolution urban Weather Research and Forecasting (WRF) simulations resolve interactions among urban form, near-surface meteorology, atmospheric dynamics, and building-energy processes, but their computational cost limits repeated scenario analysis. This study compares eight pointwise temporal surrogates (Persistence, previous-day same-hour, Ridge, Random Forest, XGBoost, multilayer perceptron (MLP), gated recurrent unit (GRU), and long short-term memory (LSTM)) for recursive 24 h prediction of T2, Q2, W10, and WRF-BEM diagnostic air-conditioning power density (AC) from a 120 h WRF-BEP/BEM-LCZ heatwave simulation over Chongqing. Model development used three expanding-window folds, fold-specific scaling, fixed monitoring cells, and five fixed seeds for stochastic families; the h97–h120 full-domain test remained untouched until all modeling choices had been finalized. Station comparisons provided a limited check on the physical plausibility of the WRF meteorological fields; AC was not validated against metered energy use. Model performance depended strongly on the target variable. Random Forest performed best for T2 (1.081 ± 0.002 °C), LSTM for Q2 (0.969 ± 0.063 g/kg), and MLP for AC (0.505 ± 0.187 W/m2), whereas deterministic Ridge was best for W10 (1.095 m/s). W10 had weaker temporal memory (median lag-1/lag-24 ACF 0.581/0.211). Ridge outperformed the Previous-day baseline for 75.0% of grid cells, although errors increased sharply in the top wind-speed decile and during rapid changes. For this single-event, single-city benchmark, recurrent models did not improve W10 accuracy over Ridge. Model choice depended on the target and the balance between accuracy, spatial fidelity, and computational cost; transfer to other events and meteorological regimes remains untested.

LandVol. 15(9)
Lawrence Livermore National Laboratory (US), Chongqing University (CN), The University of Queensland (AU), Chongqing Jiaotong University (CN)
Sustainable cities and communities
Openalex Percentile: Top 18%
Urban Heat Island Mitigation
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.