Benchmarking Deep Learning Against Statistical Baselines and a Physical Climate-Model Comparator for Station-Scale Meteorological Forecasting: A 100-Station Study from the Western Balkans

Benchmarking deep learning forecasters against classical and physically based numerical baselines remains uncommon in the time-series forecasting literature. Meteorological station networks offer an under-exploited evaluation environment, uniquely providing a physically based climate-model comparator alongside standard baselines. We evaluated eight forecasting approaches—climatology, SARIMA, Random Forest, and five deep learning architectures (TFT, N-HiTS, PatchTST, TiDE, xLSTM)—against bias-corrected output from a five-member CMIP6 ensemble, on 100 meteorological stations across four Western Balkan countries (monthly temperature and precipitation, 1961–2020), using non-parametric significance testing, a rolling-origin backtest (five windows, 2011–2020), and a five-seed robustness check. For temperature, all five deep learning architectures achieved lower MAE than the classical baselines (p < 10−99), though PatchTST’s advantage over climatology was not significant; the best-performing architecture varied across seeds and evaluation windows, so we characterise a leading cluster (N-HiTS, TFT, TiDE, PatchTST) rather than a single winner. The primary temperature advantage was geographically broad-based, while the comparison against the physical-model baseline was robust to the choice of comparator GCM. For precipitation, by contrast, a simple climatological-mean baseline outperformed all five deep learning architectures with no exception across all five rolling-origin windows. The deep learning advantage over classical and physical baselines is thus variable-specific rather than universal. Meteorological station networks, combined with a physically based climate-model comparator, constitute a well-suited evaluation environment for the broader time series forecasting community.

Authors

Institutions

Publication Details

Journal
AI
Published
2026-08-26
DOI
https://doi.org/10.3390/ai7090329
Primary Topic
Climate variability and models
Type
article
Field-Weighted Citation Impact
0.00

Funders

Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Benchmarking Deep Learning Against Statistical Baselines and a Physical Climate-Model Comparator for Station-Scale Meteorological Forecasting: A 100-Station Study from the Western Balkans

Mlađen Jovanović, Rastislav Stojsavljević, Ivica Djalović, Dalibor Nikolić et al.
AI
Climate variability and models
article

Benchmarking Deep Learning Against Statistical Baselines and a Physical Climate-Model Comparator for Station-Scale Meteorological Forecasting: A 100-Station Study from the Western Balkans

Mlađen Jovanović, Rastislav Stojsavljević, Ivica Djalović, Dalibor Nikolić, Dejan Stojanović, Ivan Vitezović, Sara Pavkov
article en

Abstract

Benchmarking deep learning forecasters against classical and physically based numerical baselines remains uncommon in the time-series forecasting literature. Meteorological station networks offer an under-exploited evaluation environment, uniquely providing a physically based climate-model comparator alongside standard baselines. We evaluated eight forecasting approaches—climatology, SARIMA, Random Forest, and five deep learning architectures (TFT, N-HiTS, PatchTST, TiDE, xLSTM)—against bias-corrected output from a five-member CMIP6 ensemble, on 100 meteorological stations across four Western Balkan countries (monthly temperature and precipitation, 1961–2020), using non-parametric significance testing, a rolling-origin backtest (five windows, 2011–2020), and a five-seed robustness check. For temperature, all five deep learning architectures achieved lower MAE than the classical baselines (p < 10−99), though PatchTST’s advantage over climatology was not significant; the best-performing architecture varied across seeds and evaluation windows, so we characterise a leading cluster (N-HiTS, TFT, TiDE, PatchTST) rather than a single winner. The primary temperature advantage was geographically broad-based, while the comparison against the physical-model baseline was robust to the choice of comparator GCM. For precipitation, by contrast, a simple climatological-mean baseline outperformed all five deep learning architectures with no exception across all five rolling-origin windows. The deep learning advantage over classical and physical baselines is thus variable-specific rather than universal. Meteorological station networks, combined with a physically based climate-model comparator, constitute a well-suited evaluation environment for the broader time series forecasting community.

AIVol. 7(9)
University of Kragujevac (RS), University of Novi Sad (RS), Institute of Field and Vegetable Crops (RS), Institute of Lowland Forestry and Environment (RS)
Provincial Secretariat for Science and Technological Development
Climate action
Openalex Percentile: Top 13%
Climate variability and models
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.