Climate-stratified assessment of PM2.5 using machine learning: Geographic controls dominate meteorological factors

Fine particulate matter (PM2.5) is a major environmental concern, yet pollution assessment models are rarely tested across fundamentally different climate regimes. We develop a climate-stratified machine-learning framework using 49,997 monthly MERRA-2 satellite-assimilated samples (2019-2023) across five latitude-band climate zones, evaluated through Leave-One-Climate-Out (LOCO) cross-validation with nested hyperparameter tuning to prevent data leakage. Tuned LightGBM achieves a global LOCO R2 = 0.367 (MAE =6.43 μg m-3), but predictability varies substantially by zone: subtropical regions reach R2 = 0.604, while tropical regions achieve only R2 = 0.142, reflecting fundamental climate-dependent limits on meteorology-only models. SHAP analysis shows geographic coordinates-acting as proxies for emission patterns and climatological transport pathways-account for 64% of predictive importance on average across zones, compared to 23% for meteorological variables, indicating that PM2.5 spatial structure is driven more by where emissions occur than by month-to-month weather variability. To verify that coordinates add genuine predictive value beyond spatial autocorrelation, a retrain ablation confirms that removing latitude and longitude collapses model performance across all zones (mean R2 drops from 0.367 to -0.099), while an IDW spatial baseline confirms that the ML model captures meteorological signal beyond simple proximity interpolation (+98% R2 improvement over IDW). Temporal validation (2019-2021 training, 2022-2023 testing) yields R2 = 0.544, confirming LOCO is the more demanding test. Ensemble stacking degrades performance relative to the zone-tuned model, providing evidence of feature-set saturation rather than model weakness. This reproducible, open-source framework offers a useful benchmark for reanalysis-based PM2.5 assessment, particularly in data-sparse regions where ground monitoring is limited.

Authors

Institutions

Publication Details

Journal
PLoS ONE
Published
2026-09-11
DOI
https://doi.org/10.1371/journal.pone.0347591
Primary Topic
Atmospheric aerosols and clouds
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Climate-stratified assessment of PM2.5 using machine learning: Geographic controls dominate meteorological factors

Ruhaidah Samsudin, Someyo kamal Utsho, Md Rawjatul Fahim
PLoS ONE
Atmospheric aerosols and clouds
article

Climate-stratified assessment of PM2.5 using machine learning: Geographic controls dominate meteorological factors

Ruhaidah Samsudin, Someyo kamal Utsho, Md Rawjatul Fahim
article en

Abstract

Fine particulate matter (PM2.5) is a major environmental concern, yet pollution assessment models are rarely tested across fundamentally different climate regimes. We develop a climate-stratified machine-learning framework using 49,997 monthly MERRA-2 satellite-assimilated samples (2019-2023) across five latitude-band climate zones, evaluated through Leave-One-Climate-Out (LOCO) cross-validation with nested hyperparameter tuning to prevent data leakage. Tuned LightGBM achieves a global LOCO R2 = 0.367 (MAE =6.43 μg m-3), but predictability varies substantially by zone: subtropical regions reach R2 = 0.604, while tropical regions achieve only R2 = 0.142, reflecting fundamental climate-dependent limits on meteorology-only models. SHAP analysis shows geographic coordinates-acting as proxies for emission patterns and climatological transport pathways-account for 64% of predictive importance on average across zones, compared to 23% for meteorological variables, indicating that PM2.5 spatial structure is driven more by where emissions occur than by month-to-month weather variability. To verify that coordinates add genuine predictive value beyond spatial autocorrelation, a retrain ablation confirms that removing latitude and longitude collapses model performance across all zones (mean R2 drops from 0.367 to -0.099), while an IDW spatial baseline confirms that the ML model captures meteorological signal beyond simple proximity interpolation (+98% R2 improvement over IDW). Temporal validation (2019-2021 training, 2022-2023 testing) yields R2 = 0.544, confirming LOCO is the more demanding test. Ensemble stacking degrades performance relative to the zone-tuned model, providing evidence of feature-set saturation rather than model weakness. This reproducible, open-source framework offers a useful benchmark for reanalysis-based PM2.5 assessment, particularly in data-sparse regions where ground monitoring is limited.

PLoS ONEVol. 21(9)
Daffodil International University (BD), University of Technology Malaysia (MY)
Climate action
Openalex Percentile: Top 14%
Atmospheric aerosols and clouds
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.