Machine Learning Classification of Elevated Discharge in the Koksu River Basin, Kazakhstan: Benchmarking Against Persistence and Illustrative DEM-Based Inundation Scenarios

In data-sparse Central Asian watersheds, machine learning and Geographic Information Systems (GIS) are increasingly applied to hydrological hazard classification. This article compares Random Forest (RF), XGBoost, and Long Short-Term Memory (LSTM) models for daily elevated-discharge classification in the Koksu River basin, Zhetysu Region, Kazakhstan (1614 km2, semi-arid, snowmelt- and glacier-influenced climate; 2005–2023), using verified discharge, precipitation, and temperature records from four monitoring stations. The elevated-discharge threshold (80th percentile, ~90.5 m3/s) was computed using only the training period (2005–2017) to avoid temporal data leakage into validation and test periods. Class imbalance, addressed via class weighting, was moderate under this threshold. All models were benchmarked against a naive persistence baseline. Contrary to expectation, the persistence baseline achieved the highest Critical Success Index (CSI = 0.821, 95% block-bootstrap CI [0.725, 0.891]), narrowly ahead of XGBoost (CSI = 0.813, [0.703, 0.897]), Random Forest (CSI = 0.809, [0.691, 0.897]), and LSTM (CSI = 0.763, [0.630, 0.873]); the confidence intervals overlap substantially, indicating no statistically distinguishable advantage of any trained model over simple persistence in this basin. Feature-importance and ablation analyses confirmed that lagged discharge, not precipitation, drove nearly all predictive skill: a discharge-lags-only model performed as well as or better than the complete pipeline, while a meteorology-and-seasonality-only model performed markedly worse (CSI ≈ 0.57–0.59). As a preliminary, exploratory illustration rather than a core result, Gumbel-derived return-period discharges (Q10, Q25, and Q100) were separately translated into illustrative, uncalibrated, DEM-proximity-based inundation extents, a candidate direction for future work rather than an engineering-grade flood-mapping contribution of this study. These results indicate that, in this basin and period, machine learning classifiers provide no clearly demonstrated advantage over simple discharge persistence, underscoring the necessity of routine persistence benchmarking before adopting machine learning approaches for early warning in Central Asian watersheds facing increasing flood risk under climate change.

Authors

Institutions

Publication Details

Journal
Water
Published
2026-09-21
DOI
https://doi.org/10.3390/w18182352
Primary Topic
Flood Risk Assessment and Management
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Machine Learning Classification of Elevated Discharge in the Koksu River Basin, Kazakhstan: Benchmarking Against Persistence and Illustrative DEM-Based Inundation Scenarios

Sangchul Lee, Asset Arystanov, A.N. Munaitpasova, Jay Sagin et al.
Water
Flood Risk Assessment and Management
article

Machine Learning Classification of Elevated Discharge in the Koksu River Basin, Kazakhstan: Benchmarking Against Persistence and Illustrative DEM-Based Inundation Scenarios

Sangchul Lee, Asset Arystanov, A.N. Munaitpasova, Jay Sagin, Ranida Arystanova, О.С. Курманбаев, Abzal Kalygulov, R. V. Yussupov, Talgat Usmanov, Sholpan Kulbekova
article en

Abstract

In data-sparse Central Asian watersheds, machine learning and Geographic Information Systems (GIS) are increasingly applied to hydrological hazard classification. This article compares Random Forest (RF), XGBoost, and Long Short-Term Memory (LSTM) models for daily elevated-discharge classification in the Koksu River basin, Zhetysu Region, Kazakhstan (1614 km2, semi-arid, snowmelt- and glacier-influenced climate; 2005–2023), using verified discharge, precipitation, and temperature records from four monitoring stations. The elevated-discharge threshold (80th percentile, ~90.5 m3/s) was computed using only the training period (2005–2017) to avoid temporal data leakage into validation and test periods. Class imbalance, addressed via class weighting, was moderate under this threshold. All models were benchmarked against a naive persistence baseline. Contrary to expectation, the persistence baseline achieved the highest Critical Success Index (CSI = 0.821, 95% block-bootstrap CI [0.725, 0.891]), narrowly ahead of XGBoost (CSI = 0.813, [0.703, 0.897]), Random Forest (CSI = 0.809, [0.691, 0.897]), and LSTM (CSI = 0.763, [0.630, 0.873]); the confidence intervals overlap substantially, indicating no statistically distinguishable advantage of any trained model over simple persistence in this basin. Feature-importance and ablation analyses confirmed that lagged discharge, not precipitation, drove nearly all predictive skill: a discharge-lags-only model performed as well as or better than the complete pipeline, while a meteorology-and-seasonality-only model performed markedly worse (CSI ≈ 0.57–0.59). As a preliminary, exploratory illustration rather than a core result, Gumbel-derived return-period discharges (Q10, Q25, and Q100) were separately translated into illustrative, uncalibrated, DEM-proximity-based inundation extents, a candidate direction for future work rather than an engineering-grade flood-mapping contribution of this study. These results indicate that, in this basin and period, machine learning classifiers provide no clearly demonstrated advantage over simple discharge persistence, underscoring the necessity of routine persistence benchmarking before adopting machine learning approaches for early warning in Central Asian watersheds facing increasing flood risk under climate change.

WaterVol. 18(18)
Western Michigan University (US), Kazakh-British Technical University (KZ), Al-Farabi Kazakh National University (KZ), Korea University (KR), Satbayev University (KZ)
Climate action
Openalex Percentile: Top 14%
Flood Risk Assessment and Management
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.