Forecasting of PM2.5/PM10 Using Machine Learning: A Benchmarking Study Based on Open Air Quality IoT Datasets

Particulate matter (PM2.5 and PM10) forecasting is increasingly framed as an applied machine-learning problem operating on real-world environmental sensor infrastructure, yet most benchmarking studies evaluate models on a single, curated station rather than on the heterogeneous, imperfect data that operational networks actually produce, and on a single chronological test split whose representativeness is rarely questioned. This study benchmarks four supervised model families, Random Forest (RF), Support Vector Regression with an RBF kernel (SVM), a feed-forward Neural Network (NN), and a Long Short-Term Memory network (LSTM), against a naïve persistence baseline and a conventional autoregressive baseline (SARIMA), across five of six monitoring stations of the Greek National Air Pollution Monitoring Network (EDPAR), a sixth, short-record rural station retained for exploratory analysis only, selected to span contrasting emission regimes and record lengths of 4 to 24 years. Using a univariate, past-only feature set (autoregressive lags, trailing rolling statistics, and cyclical calendar encodings), we forecast both next-day concentration and next-week maximum concentration for PM2.5 and PM10 independently at each station, under a two-stage evaluation design: a 40-window sliding model-selection stage, which selects each model family’s hyperparameters by mean R2 across many chronological cutoffs, and a single held-out 20% model-assessment stage on data never used for selection. SVM is the most broadly reliable model family under both stages, winning 11 of 20 station/pollutant/horizon combinations under model selection and 13 of 20 under model assessment, with its advantage most pronounced at the seven-day-maximum horizon; RF, NN and LSTM are each competitive at specific stations, but none is reliably dominant. The two evaluation stages agree on the winning model in only 10 of 20 combinations, illustrating that model rankings from a single chronological split can depend materially on the specific evaluation window chosen. SARIMA underperforms the best machine-learning model in all 20 of 20 combinations and underperforms naïve persistence itself at several stations, indicating that the machine-learning models’ advantage reflects genuine predictive skill rather than merely the seasonal and autoregressive information already available to any conventional time-series method. Achievable R2 is generally, though not universally, lower for PM10 than for PM2.5, reflecting PM10’s larger coarse-mode, episodically-driven component; station-specific exceptions to this pattern are attributable to long-term trend and test-set variance-compression effects rather than to any intrinsic reversal of pollutant forecastability. These results indicate that model selection for PM forecasting should be validated across multiple evaluation windows rather than a single split, that conventional statistical baselines remain a necessary comparison point, and that reported gains from more complex architectures should be interpreted cautiously in light of test-set non-stationarity and the sensitivity of model rankings to the choice of evaluation window.

Authors

Institutions

Publication Details

Journal
Electronics
Published
2026-10-01
DOI
https://doi.org/10.3390/electronics15194496
Primary Topic
Air Quality Monitoring and Forecasting
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Forecasting of PM2.5/PM10 Using Machine Learning: A Benchmarking Study Based on Open Air Quality IoT Datasets

Yiannis Kiouvrekis, Angeliki I. Katsafadou, Theodor Panagiotakopoulos, Christos E. Christakis et al.
Electronics
Air Quality Monitoring and Forecasting
article

Forecasting of PM2.5/PM10 Using Machine Learning: A Benchmarking Study Based on Open Air Quality IoT Datasets

Yiannis Kiouvrekis, Angeliki I. Katsafadou, Theodor Panagiotakopoulos, Christos E. Christakis, Ioannis Psomadakis, Christina L. Metallidou
article en

Abstract

Particulate matter (PM2.5 and PM10) forecasting is increasingly framed as an applied machine-learning problem operating on real-world environmental sensor infrastructure, yet most benchmarking studies evaluate models on a single, curated station rather than on the heterogeneous, imperfect data that operational networks actually produce, and on a single chronological test split whose representativeness is rarely questioned. This study benchmarks four supervised model families, Random Forest (RF), Support Vector Regression with an RBF kernel (SVM), a feed-forward Neural Network (NN), and a Long Short-Term Memory network (LSTM), against a naïve persistence baseline and a conventional autoregressive baseline (SARIMA), across five of six monitoring stations of the Greek National Air Pollution Monitoring Network (EDPAR), a sixth, short-record rural station retained for exploratory analysis only, selected to span contrasting emission regimes and record lengths of 4 to 24 years. Using a univariate, past-only feature set (autoregressive lags, trailing rolling statistics, and cyclical calendar encodings), we forecast both next-day concentration and next-week maximum concentration for PM2.5 and PM10 independently at each station, under a two-stage evaluation design: a 40-window sliding model-selection stage, which selects each model family’s hyperparameters by mean R2 across many chronological cutoffs, and a single held-out 20% model-assessment stage on data never used for selection. SVM is the most broadly reliable model family under both stages, winning 11 of 20 station/pollutant/horizon combinations under model selection and 13 of 20 under model assessment, with its advantage most pronounced at the seven-day-maximum horizon; RF, NN and LSTM are each competitive at specific stations, but none is reliably dominant. The two evaluation stages agree on the winning model in only 10 of 20 combinations, illustrating that model rankings from a single chronological split can depend materially on the specific evaluation window chosen. SARIMA underperforms the best machine-learning model in all 20 of 20 combinations and underperforms naïve persistence itself at several stations, indicating that the machine-learning models’ advantage reflects genuine predictive skill rather than merely the seasonal and autoregressive information already available to any conventional time-series method. Achievable R2 is generally, though not universally, lower for PM10 than for PM2.5, reflecting PM10’s larger coarse-mode, episodically-driven component; station-specific exceptions to this pattern are attributable to long-term trend and test-set variance-compression effects rather than to any intrinsic reversal of pollutant forecastability. These results indicate that model selection for PM forecasting should be validated across multiple evaluation windows rather than a single split, that conventional statistical baselines remain a necessary comparison point, and that reported gains from more complex architectures should be interpreted cautiously in light of test-set non-stationarity and the sensitivity of model rankings to the choice of evaluation window.

ElectronicsVol. 15(19)
University of Thessaly (GR), University of Nicosia (CY), University of Patras (GR), Technological Educational Institute of Thessaly (GR)
Industry, innovation and infrastructure
Openalex Percentile: Top 19%
Air Quality Monitoring and Forecasting
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.