Ensemble machine learning-based comparative prediction of NO2 and O3 concentrations in Bengaluru, India

Accurate prediction of urban air pollutants is essential for effective air quality management and public health protection. Although machine learning (ML) models have been widely applied to air quality prediction, limited attention has been given to systematic comparative assessment of pollutants with distinct atmospheric formation mechanisms within a unified modeling framework. In particular, comparative evidence integrating model benchmarking, multi-metric evaluation, model generalization, and interpretability for both primary and secondary pollutants remain limited. This study systematically benchmarked seven ensemble ML algorithms—Random Forest, Bagging, Extra Trees, XGBoost, LightGBM, CatBoost, and AdaBoost—for predicting nitrogen dioxide (NO 2 ), a primary pollutant, and ozone (O 3 ), a secondary pollutant, using 7,627 daily observations collected from six Continuous Ambient Air Quality Monitoring Stations (CAAQMS) in Bengaluru, India, between 2021 and 2024. Model performance was evaluated using randomized hyperparameter optimization, five-fold cross-validation, and multiple performance metrics, including coefficient of determination (R 2 ), root mean square error (RMSE), mean absolute error (MAE), and symmetric mean absolute percentage error (sMAPE). CatBoost achieved the highest predictive performance for both NO 2 (R 2 = 0.684, RMSE = 7.48, MAE = 4.84) and O 3 (R 2 = 0.646, RMSE = 7.40, MAE = 4.85) on the independent test dataset while maintaining the most favourable balance between predictive accuracy and model generalization. The comparatively lower predictive performance for O 3 reflects the greater complexity of its nonlinear photochemical formation relative to the emission-driven behaviour of NO 2 . Furthermore, SHAP-based explainable artificial intelligence (XAI) analysis identified NH 3 as the dominant predictor of NO 2 and NO 2 as the most influential predictor of O 3 , providing physically interpretable insights into the contrasting atmospheric processes governing the two pollutants. By integrating comparative ensemble benchmarking, pollutant-specific generalization assessment, multi-metric evaluation, and explainable ML within a common framework, this study provides a more comprehensive basis for evaluating predictive reliability and atmospheric interpretability across primary and secondary urban pollutants. The findings demonstrate the value of comparative benchmarking, comprehensive multi-metric evaluation, and explainable machine learning for developing reliable and interpretable urban air quality prediction systems.

Authors

Institutions

Publication Details

Journal
Scientific Reports
Published
2026-09-05
DOI
https://doi.org/10.1038/s41598-026-70145-w
Primary Topic
Air Quality Monitoring and Forecasting
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Ensemble machine learning-based comparative prediction of NO2 and O3 concentrations in Bengaluru, India

Avanthika Kumar, Sophia Lawrence, Preethi Baskaran, Srimuruganandam Bhathmanabhan et al.
Scientific Reports
Air Quality Monitoring and Forecasting
article

Ensemble machine learning-based comparative prediction of NO2 and O3 concentrations in Bengaluru, India

Avanthika Kumar, Sophia Lawrence, Preethi Baskaran, Srimuruganandam Bhathmanabhan, Tanya Suresh
article en

Abstract

Accurate prediction of urban air pollutants is essential for effective air quality management and public health protection. Although machine learning (ML) models have been widely applied to air quality prediction, limited attention has been given to systematic comparative assessment of pollutants with distinct atmospheric formation mechanisms within a unified modeling framework. In particular, comparative evidence integrating model benchmarking, multi-metric evaluation, model generalization, and interpretability for both primary and secondary pollutants remain limited. This study systematically benchmarked seven ensemble ML algorithms—Random Forest, Bagging, Extra Trees, XGBoost, LightGBM, CatBoost, and AdaBoost—for predicting nitrogen dioxide (NO 2 ), a primary pollutant, and ozone (O 3 ), a secondary pollutant, using 7,627 daily observations collected from six Continuous Ambient Air Quality Monitoring Stations (CAAQMS) in Bengaluru, India, between 2021 and 2024. Model performance was evaluated using randomized hyperparameter optimization, five-fold cross-validation, and multiple performance metrics, including coefficient of determination (R 2 ), root mean square error (RMSE), mean absolute error (MAE), and symmetric mean absolute percentage error (sMAPE). CatBoost achieved the highest predictive performance for both NO 2 (R 2 = 0.684, RMSE = 7.48, MAE = 4.84) and O 3 (R 2 = 0.646, RMSE = 7.40, MAE = 4.85) on the independent test dataset while maintaining the most favourable balance between predictive accuracy and model generalization. The comparatively lower predictive performance for O 3 reflects the greater complexity of its nonlinear photochemical formation relative to the emission-driven behaviour of NO 2 . Furthermore, SHAP-based explainable artificial intelligence (XAI) analysis identified NH 3 as the dominant predictor of NO 2 and NO 2 as the most influential predictor of O 3 , providing physically interpretable insights into the contrasting atmospheric processes governing the two pollutants. By integrating comparative ensemble benchmarking, pollutant-specific generalization assessment, multi-metric evaluation, and explainable ML within a common framework, this study provides a more comprehensive basis for evaluating predictive reliability and atmospheric interpretability across primary and secondary urban pollutants. The findings demonstrate the value of comparative benchmarking, comprehensive multi-metric evaluation, and explainable machine learning for developing reliable and interpretable urban air quality prediction systems.

Scientific Reports
Vellore Institute of Technology University (IN)
Sustainable cities and communities
Openalex Percentile: Top 17%
Air Quality Monitoring and Forecasting
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.