Expert rules versus machine learning for lysosomal storage disorder identification: a comparative analysis using clinical data
Abstract Early diagnosis of rare lysosomal storage disorders, such as Gaucher disease (GD), acid sphingomyelinase deficiency (ASMD), and Saposin C deficiency, is critical but often delayed. Electronic health records (EHRs) provide opportunities for screening, but class imbalance challenges computational methods. The real-world performance of one-class classification (OCC) machine learning (ML) models versus clinician-derived expert rules remains poorly understood. We compare a clinician-curated expert rule system, based on phenotypes like organomegaly and cytopenias, against eight OCC models on 572 patients (13 GD, 8 ASMD, 3 Saposin C, 548 controls) from a United Arab Emirates health system. The evaluation used imbalance-aware metrics including recall, precision, F1- and F2-scores, balanced accuracy, the area under the precision–recall curve, and the number needed to screen (NNS). The expert rule system demonstrated high, stable recall (76.9%–100.0%) with a clinically feasible workload (NNS: 3.0–7.7). In a matched comparison, rules consistently outperformed the best OCC models (F1-score: GD 0.749 vs 0.626; ASMD 0.751 vs 0.580; Saposin 0.631 vs 0.360). ML models exhibited severe instability as rarity increased, with catastrophic failures (0% precision) on the ultra-rare Saposin cohort. Across all three disorders in this EHR cohort, expert rules delivered more stable recall and a more clinically feasible screening workload (NNS) than one-class ML as disease rarity increased. This stability proved critical, as the primary clinical objective in rare disease screening is to ensure no cases are missed (i.e., to maximize recall). The expert rules provided this dependable, high-recall screen, whereas selected OCC models that raised precision did so at an unacceptable cost of reduced sensitivity and stability. Given the ultra-rare Saposin cohort, OCC methods were unreliable in our dataset, whereas expert rules preserved sensitivity. Overall, our results motivate a hybrid strategy in which expert rules set a reliable recall floor and ML is used only as a secondary tool to re-rank or triage within the flagged subset.
Authors
- Amal Al Tenaiji
- Jaloliddin Rustamov (ORCID: https://orcid.org/0000-0003-3701-1611)
- Hany Alashwal (ORCID: https://orcid.org/0000-0002-5721-5104)
- Zahiriddin Rustamov (ORCID: https://orcid.org/0000-0003-4977-1781)
- Fatma Al Jasmi
- Ayisha Manzoor (ORCID: https://orcid.org/0000-0002-3791-9892)
- Muhammad Jalal Khan (ORCID: https://orcid.org/0000-0002-6230-1760)
- Mohd Saberi Mohamad (ORCID: https://orcid.org/0000-0002-1079-4559)
- Mariam A. Alharbi (ORCID: https://orcid.org/0000-0003-1442-0674)
- Hamzah Dabool
- Nazar Zaki
- Faizah Hashimy
Publication Details
- Journal
- Scientific Reports
- Published
- 2026-09-22
- DOI
- https://doi.org/10.1038/s41598-026-71844-0
- Primary Topic
- Lysosomal Storage Disorders Research
- Type
- article
- Field-Weighted Citation Impact
- 0.00