Machine Learning Surrogate Modeling in R for Rapid Screening of Green Infrastructure Hydrological Performance in Urban Stormwater Management: A Proof-of-Concept Study Using Synthetic Data

Background: Physically-based, coupled hydrological–low-impact-development (LID) models, such as the U.S. EPA Storm Water Management Model (SWMM), estimate green infrastructure (GI) performance in detail but are computationally expensive to run across many catchment, storm, and typology combinations. Methods: This methodological proof-of-concept develops an open-source R workflow (randomForest, xgboost, caret) on a synthetic dataset of 2000 catchment–storm–typology scenarios generated from prescribed non-linear equations of imperviousness, storm return period, and the coverage of four GI typologies. None of the scenarios were generated by SWMM-LID simulation or field monitoring. Random forest, XGBoost, and a linear baseline were trained to predict synthetic peak-flow attenuation and suspended-solid (TSS) removal. Results: On the held-out synthetic test set, XGBoost and random forest reached R2 values of 0.96 and 0.92 for peak-flow attenuation (linear baseline: 0.90) and 0.94 and 0.89 for TSS reduction (linear baseline: 0.79). These values show that the models learned the synthetic response surface; they do not measure predictive skill for real GI systems. Feature importance reproduced the typology weighting embedded in the data-generating equations. Surrogate inference took about 2–4 ms (measured), compared with tens of minutes typically reported in the literature for SWMM-LID runs (not measured here); this is an indicative comparison, not a controlled benchmark. Conclusions: This study is a methodological demonstration only and is not a validated hydrological surrogate. Real application requires retraining and validation using SWMM-LID simulation ensembles or field-monitored data.

Authors

Institutions

Publication Details

Journal
Water
Published
2026-09-28
DOI
https://doi.org/10.3390/w18192413
Primary Topic
Urban Stormwater Management Solutions
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Machine Learning Surrogate Modeling in R for Rapid Screening of Green Infrastructure Hydrological Performance in Urban Stormwater Management: A Proof-of-Concept Study Using Synthetic Data

Ivona Škultétyová, Ján Ilavský, Raghad A. Awad, Štefan Stanko et al.
Water
Urban Stormwater Management Solutions
article

Machine Learning Surrogate Modeling in R for Rapid Screening of Green Infrastructure Hydrological Performance in Urban Stormwater Management: A Proof-of-Concept Study Using Synthetic Data

Ivona Škultétyová, Ján Ilavský, Raghad A. Awad, Štefan Stanko, Danka Barloková
article en

Abstract

Background: Physically-based, coupled hydrological–low-impact-development (LID) models, such as the U.S. EPA Storm Water Management Model (SWMM), estimate green infrastructure (GI) performance in detail but are computationally expensive to run across many catchment, storm, and typology combinations. Methods: This methodological proof-of-concept develops an open-source R workflow (randomForest, xgboost, caret) on a synthetic dataset of 2000 catchment–storm–typology scenarios generated from prescribed non-linear equations of imperviousness, storm return period, and the coverage of four GI typologies. None of the scenarios were generated by SWMM-LID simulation or field monitoring. Random forest, XGBoost, and a linear baseline were trained to predict synthetic peak-flow attenuation and suspended-solid (TSS) removal. Results: On the held-out synthetic test set, XGBoost and random forest reached R2 values of 0.96 and 0.92 for peak-flow attenuation (linear baseline: 0.90) and 0.94 and 0.89 for TSS reduction (linear baseline: 0.79). These values show that the models learned the synthetic response surface; they do not measure predictive skill for real GI systems. Feature importance reproduced the typology weighting embedded in the data-generating equations. Surrogate inference took about 2–4 ms (measured), compared with tens of minutes typically reported in the literature for SWMM-LID runs (not measured here); this is an indicative comparison, not a controlled benchmark. Conclusions: This study is a methodological demonstration only and is not a validated hydrological surrogate. Real application requires retraining and validation using SWMM-LID simulation ensembles or field-monitored data.

WaterVol. 18(19)
Slovak University of Technology in Bratislava (SK)
Industry, innovation and infrastructure
Openalex Percentile: Top 19%
Urban Stormwater Management Solutions
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Machine Learning Surrogate Modeling in R for Rapid Screening of Green Infrastructure Hydrological Performance in Urban Stormwater Management: A Proof-of-Concept Study Using Synthetic Data — Ivona Škultétyová, Ján Ilavský, et al. · Water (2026) | TGRS Research Map | TGRS