Rainfall frequency and uncertainty analysis in arid regions using machine learning and statistical distribution models based on a 91-year rainfall record

Abstract Accurate estimation of rainfall depths associated with standard return periods is fundamental for hydraulic design, flood risk assessment, and water resources management, particularly in arid regions. Statistical extreme value distributions have long been the standard approach for rainfall frequency analysis, whereas machine learning (ML) has recently emerged as a potential alternative. This study presents a comprehensive comparison between twenty-five statistical distribution models and three supervised ML regressors—Support Vector Machine (SVM), Random Forest (RF), and Extreme Gradient Boosting Regressor (XGBR)—for Annual Maximum Daily Rainfall (AMDR) frequency analysis using a 91-year record (1931–2021) from the Suez meteorological station, Egypt. Prior to frequency analysis, the AMDR series was verified to satisfy the assumptions of stationarity and serial independence using the Mann–Kendall test, Sen's slope estimator, and Lag-1 autocorrelation analysis. Statistical distribution models were implemented in HEC-SSP using Weibull plotting positions, whereas the ML models were developed in Python using exceedance probability as the sole predictor to enable a direct and consistent comparison with statistical frequency models. Model performance was evaluated using the coefficient of determination (R 2 ), mean squared error (MSE), root mean squared error (RMSE), and mean absolute error (MAE), while predictive uncertainty was quantified using 90% bootstrap confidence intervals. The results demonstrate that statistical distribution models consistently outperformed the ML regressors across all performance metrics. The Generalized Extreme Value distribution fitted using Maximum Likelihood Estimation (GEVMLE) achieved the highest predictive performance (R 2 = 0.968; RMSE = 1.512 mm), followed by GEVLM and GLLM. In contrast, the ML models exhibited substantially lower predictive accuracy (R 2 = 0.280 for RF, 0.258 for XGBR, and 0.087 for SVM) and showed limited capability for extrapolating rainfall extremes beyond the observed data range at longer return periods (T ≥ 25 years). The ML models also produced comparatively narrow confidence intervals, reflecting the limited variability represented within the training domain, whereas the GEV-based statistical models provided more realistic uncertainty estimates for extreme rainfall quantiles. Overall, the findings indicate that extreme value distributions, particularly the GEV models, remain more robust than ML regressors for rainfall frequency analysis when only long-term single-variable rainfall records are available. GEVMLE is therefore recommended for rainfall frequency analysis at the Suez station and similar arid coastal environments. Future research should investigate hybrid frameworks that integrate extreme value theory with machine learning using multi-variable climatic predictors to improve extrapolation capability and uncertainty representation.

Authors

Institutions

Publication Details

Journal
Scientific Reports
Published
2026-09-15
DOI
https://doi.org/10.1038/s41598-026-68681-6
Primary Topic
Hydrology and Drought Analysis
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Rainfall frequency and uncertainty analysis in arid regions using machine learning and statistical distribution models based on a 91-year rainfall record

Mohammed Hagage
Scientific Reports
Hydrology and Drought Analysis
article

Rainfall frequency and uncertainty analysis in arid regions using machine learning and statistical distribution models based on a 91-year rainfall record

Mohammed Hagage
article en

Abstract

Abstract Accurate estimation of rainfall depths associated with standard return periods is fundamental for hydraulic design, flood risk assessment, and water resources management, particularly in arid regions. Statistical extreme value distributions have long been the standard approach for rainfall frequency analysis, whereas machine learning (ML) has recently emerged as a potential alternative. This study presents a comprehensive comparison between twenty-five statistical distribution models and three supervised ML regressors—Support Vector Machine (SVM), Random Forest (RF), and Extreme Gradient Boosting Regressor (XGBR)—for Annual Maximum Daily Rainfall (AMDR) frequency analysis using a 91-year record (1931–2021) from the Suez meteorological station, Egypt. Prior to frequency analysis, the AMDR series was verified to satisfy the assumptions of stationarity and serial independence using the Mann–Kendall test, Sen's slope estimator, and Lag-1 autocorrelation analysis. Statistical distribution models were implemented in HEC-SSP using Weibull plotting positions, whereas the ML models were developed in Python using exceedance probability as the sole predictor to enable a direct and consistent comparison with statistical frequency models. Model performance was evaluated using the coefficient of determination (R 2 ), mean squared error (MSE), root mean squared error (RMSE), and mean absolute error (MAE), while predictive uncertainty was quantified using 90% bootstrap confidence intervals. The results demonstrate that statistical distribution models consistently outperformed the ML regressors across all performance metrics. The Generalized Extreme Value distribution fitted using Maximum Likelihood Estimation (GEVMLE) achieved the highest predictive performance (R 2 = 0.968; RMSE = 1.512 mm), followed by GEVLM and GLLM. In contrast, the ML models exhibited substantially lower predictive accuracy (R 2 = 0.280 for RF, 0.258 for XGBR, and 0.087 for SVM) and showed limited capability for extrapolating rainfall extremes beyond the observed data range at longer return periods (T ≥ 25 years). The ML models also produced comparatively narrow confidence intervals, reflecting the limited variability represented within the training domain, whereas the GEV-based statistical models provided more realistic uncertainty estimates for extreme rainfall quantiles. Overall, the findings indicate that extreme value distributions, particularly the GEV models, remain more robust than ML regressors for rainfall frequency analysis when only long-term single-variable rainfall records are available. GEVMLE is therefore recommended for rainfall frequency analysis at the Suez station and similar arid coastal environments. Future research should investigate hybrid frameworks that integrate extreme value theory with machine learning using multi-variable climatic predictors to improve extrapolation capability and uncertainty representation.

Scientific ReportsVol. 16(1)
National Authority for Remote Sensing and Space Sciences (EG)
Clean water and sanitation
Openalex Percentile: Top 13%
Hydrology and Drought Analysis
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.