A multi-facility data quality assessment of electronic health records for trustworthy AI in HIV public health in Zambia
Artificial intelligence (AI) holds promise for HIV disease surveillance and clinical decision support in sub-Saharan Africa, yet AI reliability and fairness depend fundamentally on data quality, a dimension rarely assessed before model development. This study presents a systematic, multi-facility data quality assessment (DQA) of Zambia’s national HIV electronic health record (EHR) system, using an active machine-learning (ML) pipeline as the diagnostic lens. We applied an eight-dimension DQA framework (completeness, conformance, plausibility, feature coverage, outcome-definition sensitivity, temporal stability, cross-facility concordance, and missingness–outcome confounding) to 246,053 patient records across six Lusaka public health facilities. Three 12-month outcomes were operationalised from routine records and are labelled descriptively to reflect exactly what was measured: recorded TB treatment (TB12m), programme-recorded interruption in treatment (IIT12m), and recorded unsuppressed viral load (UVL12m). Feature-importance and SHAP (Shapley Additive Explanations) attribution patterns from a companion AI model were examined as empirical evidence of underlying data quality problems. Data pre-processing used median imputation for numeric features, most-frequent-value imputation for categorical features, standardisation and one-hot encoding, all fitted on training data only. Education-level completeness was 1.2–7.9% across facilities, household income 4.7–19.8%, and HIV enrolment stage 21.1–24.6%. Derived engagement features were computable for only 30,295 patients (12.3%) and laboratory features for 8233 (3.3%). Among patients with engagement data, 54.3% had appointment-lateness values exceeding plausible limits (maximum 45,459 days). The three outcome labels are strongly denominator-dependent: only 25,294 patients (10.3%) had at least one post-index viral-load test, of whom 14.1% had a recorded unsuppressed result, whereas coding untested patients as suppressed yields a cohort-level UVL12m prevalence of only 1.45%. Facility-level missingness rates co-varied with outcome rates (education vs. recorded viral suppression, r = − 0.865; HIV stage vs. recorded TB treatment, r = + 0.741); these are exploratory, ecological correlations across six facilities and are hypothesis-generating rather than confirmatory. The composite Data Quality Index ranged from 53.5 to 58.3%, below a pragmatic 70% screening threshold at every facility. Sociodemographic data gaps create an algorithmic-fairness risk: AI models may respond to missing records rather than genuine clinical need. We propose eight minimum data quality requirements and a reproducible, openly available DQA toolkit for trustworthy, equitable AI in resource-limited settings, aligned with SDG 9 and SDG 10.
Authors
- Jacob Mutale (ORCID: https://orcid.org/0000-0002-8914-0645)
- Innocent Chiboma (ORCID: https://orcid.org/0009-0000-5776-6443)
- Chiyaba Njovu
- Andrew Kashoka (ORCID: https://orcid.org/0000-0003-4581-7842)
- Aaron Zimba (ORCID: https://orcid.org/0000-0002-2587-106X)
- Mwansa Lumpa (ORCID: https://orcid.org/0009-0000-2306-6318)
- Joe Phiri (ORCID: https://orcid.org/0009-0002-0336-6852)
- Mulenga Chiwele
- Trevor Sinkala
Institutions
- Centre for Infectious Disease Research in Zambia (ZM)
- Ministry of Health (ZM)
- Zambia Centre for Accountancy Studies (ZM)
Publication Details
- Journal
- Discover Artificial Intelligence
- Published
- 2026-09-21
- DOI
- https://doi.org/10.1007/s44163-026-02293-x
- Primary Topic
- Electronic Health Records Systems
- Type
- article
- Field-Weighted Citation Impact
- 0.00