Failure-risk classification in water distribution networks using hydraulic modelling and machine learning

Failure-risk prioritization in water distribution systems requires information on both pipe characteristics and the hydraulic consequences of pipe isolation. This field-based case study integrates utility records, EPANET-based shutdown simulations, and supervised machine learning to classify pipe-segment records into low-, medium-, and high-risk categories. The analysis included 512 records. To avoid post-event causality and spatial memorization, the final predictor set was restricted to four physically interpretable variables available under normal operation: pipe diameter, pipe length, mean endpoint pressure, and absolute endpoint pressure difference. Pipe, street, and node identifiers, repair duration, and all shutdown-consequence variables were excluded from the predictor matrix. A majority-class baseline, Decision Tree (DT), Random Forest (RF), and Support Vector Machine (SVM) were evaluated using stratified pipe-section-level train/validation/test splitting, randomized hyperparameter search with five-fold stratified pipe-section cross-validation, class-wise metrics, ordinal mean absolute error (MAE), quadratic weighted kappa, and severe-error rate. SVM achieved the highest held-out accuracy (0.598) and macro-F1 (0.533), whereas RF showed the most favourable ordinal error profile among the evaluated models, with ordinal MAE 0.451, quadratic weighted kappa 0.385, and a severe two-level error rate of 2.0%. In a supplementary post-selection five-fold stratified pipe-section cross-validation on the complete record dataset, RF achieved the highest mean macro-F1 (0.499 ± 0.076). The results support internal consequence-based screening within the analysed network, but they do not establish transferability to other utilities. Limitations include the single-network design, the use of demand-driven EPANET simulations, and the absence of normal-operation pipe-flow data in the verified analytical dataset.

Authors

Institutions

Publication Details

Journal
Scientific Reports
Published
2026-10-06
DOI
https://doi.org/10.1038/s41598-026-74913-6
Primary Topic
Water Systems and Optimization
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

Failure-risk classification in water distribution networks using hydraulic modelling and machine learning

Katarzyna Pietrucha-Urbanik, A. Studziński
Scientific Reports
Water Systems and Optimization
article

Failure-risk classification in water distribution networks using hydraulic modelling and machine learning

Katarzyna Pietrucha-Urbanik, A. Studziński
article en

Abstract

Failure-risk prioritization in water distribution systems requires information on both pipe characteristics and the hydraulic consequences of pipe isolation. This field-based case study integrates utility records, EPANET-based shutdown simulations, and supervised machine learning to classify pipe-segment records into low-, medium-, and high-risk categories. The analysis included 512 records. To avoid post-event causality and spatial memorization, the final predictor set was restricted to four physically interpretable variables available under normal operation: pipe diameter, pipe length, mean endpoint pressure, and absolute endpoint pressure difference. Pipe, street, and node identifiers, repair duration, and all shutdown-consequence variables were excluded from the predictor matrix. A majority-class baseline, Decision Tree (DT), Random Forest (RF), and Support Vector Machine (SVM) were evaluated using stratified pipe-section-level train/validation/test splitting, randomized hyperparameter search with five-fold stratified pipe-section cross-validation, class-wise metrics, ordinal mean absolute error (MAE), quadratic weighted kappa, and severe-error rate. SVM achieved the highest held-out accuracy (0.598) and macro-F1 (0.533), whereas RF showed the most favourable ordinal error profile among the evaluated models, with ordinal MAE 0.451, quadratic weighted kappa 0.385, and a severe two-level error rate of 2.0%. In a supplementary post-selection five-fold stratified pipe-section cross-validation on the complete record dataset, RF achieved the highest mean macro-F1 (0.499 ± 0.076). The results support internal consequence-based screening within the analysed network, but they do not establish transferability to other utilities. Limitations include the single-network design, the use of demand-driven EPANET simulations, and the absence of normal-operation pipe-flow data in the verified analytical dataset.

Scientific Reports
Rzeszów University of Technology (PL)
Openalex Percentile: Top 17%
Water Systems and Optimization
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Failure-risk classification in water distribution networks using hydraulic modelling and machine learning — Katarzyna Pietrucha-Urbanik, A. Studziński · Scientific Reports (2026) | TGRS Research Map | TGRS