Expert rules versus machine learning for lysosomal storage disorder identification: a comparative analysis using clinical data

Abstract Early diagnosis of rare lysosomal storage disorders, such as Gaucher disease (GD), acid sphingomyelinase deficiency (ASMD), and Saposin C deficiency, is critical but often delayed. Electronic health records (EHRs) provide opportunities for screening, but class imbalance challenges computational methods. The real-world performance of one-class classification (OCC) machine learning (ML) models versus clinician-derived expert rules remains poorly understood. We compare a clinician-curated expert rule system, based on phenotypes like organomegaly and cytopenias, against eight OCC models on 572 patients (13 GD, 8 ASMD, 3 Saposin C, 548 controls) from a United Arab Emirates health system. The evaluation used imbalance-aware metrics including recall, precision, F1- and F2-scores, balanced accuracy, the area under the precision–recall curve, and the number needed to screen (NNS). The expert rule system demonstrated high, stable recall (76.9%–100.0%) with a clinically feasible workload (NNS: 3.0–7.7). In a matched comparison, rules consistently outperformed the best OCC models (F1-score: GD 0.749 vs 0.626; ASMD 0.751 vs 0.580; Saposin 0.631 vs 0.360). ML models exhibited severe instability as rarity increased, with catastrophic failures (0% precision) on the ultra-rare Saposin cohort. Across all three disorders in this EHR cohort, expert rules delivered more stable recall and a more clinically feasible screening workload (NNS) than one-class ML as disease rarity increased. This stability proved critical, as the primary clinical objective in rare disease screening is to ensure no cases are missed (i.e., to maximize recall). The expert rules provided this dependable, high-recall screen, whereas selected OCC models that raised precision did so at an unacceptable cost of reduced sensitivity and stability. Given the ultra-rare Saposin cohort, OCC methods were unreliable in our dataset, whereas expert rules preserved sensitivity. Overall, our results motivate a hybrid strategy in which expert rules set a reliable recall floor and ML is used only as a secondary tool to re-rank or triage within the flagged subset.

Authors

Publication Details

Journal
Scientific Reports
Published
2026-09-22
DOI
https://doi.org/10.1038/s41598-026-71844-0
Primary Topic
Lysosomal Storage Disorders Research
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Expert rules versus machine learning for lysosomal storage disorder identification: a comparative analysis using clinical data

Amal Al Tenaiji, Jaloliddin Rustamov, Hany Alashwal, Zahiriddin Rustamov et al.
Scientific Reports
Lysosomal Storage Disorders Research
article

Expert rules versus machine learning for lysosomal storage disorder identification: a comparative analysis using clinical data

Amal Al Tenaiji, Jaloliddin Rustamov, Hany Alashwal, Zahiriddin Rustamov, Fatma Al Jasmi, Ayisha Manzoor, Muhammad Jalal Khan, Mohd Saberi Mohamad, Mariam A. Alharbi, Hamzah Dabool, Nazar Zaki, Faizah Hashimy
article en

Abstract

Abstract Early diagnosis of rare lysosomal storage disorders, such as Gaucher disease (GD), acid sphingomyelinase deficiency (ASMD), and Saposin C deficiency, is critical but often delayed. Electronic health records (EHRs) provide opportunities for screening, but class imbalance challenges computational methods. The real-world performance of one-class classification (OCC) machine learning (ML) models versus clinician-derived expert rules remains poorly understood. We compare a clinician-curated expert rule system, based on phenotypes like organomegaly and cytopenias, against eight OCC models on 572 patients (13 GD, 8 ASMD, 3 Saposin C, 548 controls) from a United Arab Emirates health system. The evaluation used imbalance-aware metrics including recall, precision, F1- and F2-scores, balanced accuracy, the area under the precision–recall curve, and the number needed to screen (NNS). The expert rule system demonstrated high, stable recall (76.9%–100.0%) with a clinically feasible workload (NNS: 3.0–7.7). In a matched comparison, rules consistently outperformed the best OCC models (F1-score: GD 0.749 vs 0.626; ASMD 0.751 vs 0.580; Saposin 0.631 vs 0.360). ML models exhibited severe instability as rarity increased, with catastrophic failures (0% precision) on the ultra-rare Saposin cohort. Across all three disorders in this EHR cohort, expert rules delivered more stable recall and a more clinically feasible screening workload (NNS) than one-class ML as disease rarity increased. This stability proved critical, as the primary clinical objective in rare disease screening is to ensure no cases are missed (i.e., to maximize recall). The expert rules provided this dependable, high-recall screen, whereas selected OCC models that raised precision did so at an unacceptable cost of reduced sensitivity and stability. Given the ultra-rare Saposin cohort, OCC methods were unreliable in our dataset, whereas expert rules preserved sensitivity. Overall, our results motivate a hybrid strategy in which expert rules set a reliable recall floor and ML is used only as a secondary tool to re-rank or triage within the flagged subset.

Scientific Reports
Openalex Percentile: Top 11%
Lysosomal Storage Disorders Research
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.