Rainfall frequency and uncertainty analysis in arid regions using machine learning and statistical distribution models based on a 91-year rainfall record
Abstract Accurate estimation of rainfall depths associated with standard return periods is fundamental for hydraulic design, flood risk assessment, and water resources management, particularly in arid regions. Statistical extreme value distributions have long been the standard approach for rainfall frequency analysis, whereas machine learning (ML) has recently emerged as a potential alternative. This study presents a comprehensive comparison between twenty-five statistical distribution models and three supervised ML regressors—Support Vector Machine (SVM), Random Forest (RF), and Extreme Gradient Boosting Regressor (XGBR)—for Annual Maximum Daily Rainfall (AMDR) frequency analysis using a 91-year record (1931–2021) from the Suez meteorological station, Egypt. Prior to frequency analysis, the AMDR series was verified to satisfy the assumptions of stationarity and serial independence using the Mann–Kendall test, Sen's slope estimator, and Lag-1 autocorrelation analysis. Statistical distribution models were implemented in HEC-SSP using Weibull plotting positions, whereas the ML models were developed in Python using exceedance probability as the sole predictor to enable a direct and consistent comparison with statistical frequency models. Model performance was evaluated using the coefficient of determination (R 2 ), mean squared error (MSE), root mean squared error (RMSE), and mean absolute error (MAE), while predictive uncertainty was quantified using 90% bootstrap confidence intervals. The results demonstrate that statistical distribution models consistently outperformed the ML regressors across all performance metrics. The Generalized Extreme Value distribution fitted using Maximum Likelihood Estimation (GEVMLE) achieved the highest predictive performance (R 2 = 0.968; RMSE = 1.512 mm), followed by GEVLM and GLLM. In contrast, the ML models exhibited substantially lower predictive accuracy (R 2 = 0.280 for RF, 0.258 for XGBR, and 0.087 for SVM) and showed limited capability for extrapolating rainfall extremes beyond the observed data range at longer return periods (T ≥ 25 years). The ML models also produced comparatively narrow confidence intervals, reflecting the limited variability represented within the training domain, whereas the GEV-based statistical models provided more realistic uncertainty estimates for extreme rainfall quantiles. Overall, the findings indicate that extreme value distributions, particularly the GEV models, remain more robust than ML regressors for rainfall frequency analysis when only long-term single-variable rainfall records are available. GEVMLE is therefore recommended for rainfall frequency analysis at the Suez station and similar arid coastal environments. Future research should investigate hybrid frameworks that integrate extreme value theory with machine learning using multi-variable climatic predictors to improve extrapolation capability and uncertainty representation.
Authors
- Mohammed Hagage
Institutions
- National Authority for Remote Sensing and Space Sciences (EG)
Publication Details
- Journal
- Scientific Reports
- Published
- 2026-09-15
- DOI
- https://doi.org/10.1038/s41598-026-68681-6
- Primary Topic
- Hydrology and Drought Analysis
- Type
- article
- Field-Weighted Citation Impact
- 0.00