Evaluation of a National Health Service Machine-Learning Model for Hypertension Case-Finding: Retrospective Cohort Study

Abstract Background Hypertension is a leading preventable cause of cardiovascular disease, yet a substantial proportion of adults remain undiagnosed, limiting opportunities for early intervention. A predictive model was commissioned by the North West London (NWL) Integrated Care Board to identify undiagnosed hypertension. The model was developed using health records from the Whole Systems Integrated Care (WSIC) database. Objective We aimed to independently evaluate the predictive performance of the model as it would be encountered in deployment, how performance varied by demographic characteristics, and practical utility. Methods To evaluate the predictive model, we conducted a retrospective cohort study of 1,802,920 individuals aged 16 years or older, registered with a general practice in NWL, and with no prior diagnosis of hypertension from May 2023 to May 2024. We assessed the model’s predictions against recorded hypertension status using medical diagnoses and blood pressure records. Logistic regression models were used to assess the sensitivity and specificity of the model’s predictions by sociodemographic groups. We also compared the model’s performance against a more interpretable regression approach. Results The model yielded an overall sensitivity of 62.7% (95% CI 62.5-62.8) and specificity of 60.7% (95% CI 60.5-60.8). Positive predictive value ranged from 31.5% (95% CI 31.2-31.8) to 42.9% (95% CI 42.5-43.2), and negative predictive value ranged from 77.6% (95% CI 77.2-77.9) to 84.9% (95% CI 84.7-85.2). Sensitivity was higher in older adults and Black patients; specificity was higher in younger adults, female patients, and White patients. Overall, sensitivity was higher for those living in areas of higher socioeconomic deprivation, while specificity was lower. These effects plateaued in the 2 least deprived quintiles of deprivation, which were comparable in both sensitivity and specificity. Predictions varied by age, with 96.2% (58,951/61,281) of those aged 70 to 79 predicted to have hypertension, whereas 0.08% of those aged 20 to 39 were predicted to have the condition. The model’s performance was comparable with a more interpretable logistic regression model. Conclusions Despite the model’s relatively good performance for those without hypertension, the positive predictive value was low, and a significant proportion of true cases remained undetected. Furthermore, there was considerable variation in performance associated with demographic characteristics, suggesting tailored approaches to case-finding may be beneficial in ensuring equity across demographic groups. Especially given the importance of understanding possible biases in predictive models, we recommend that, where there is no loss in performance, more parsimonious, transparent models be selected for prediction in health care settings. The findings of this evaluation can guide the practical application of the model, inform enhancements, direct targeted screening initiatives, and support cost-benefit analyses for broader implementation to improve hypertension management.

Authors

Publication Details

Journal
Journal of Medical Internet Research
Published
2026-09-15
DOI
https://doi.org/10.2196/87084
Primary Topic
Blood Pressure and Hypertension Studies
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Evaluation of a National Health Service Machine-Learning Model for Hypertension Case-Finding: Retrospective Cohort Study

Thomas Beaney, Ahmad Alkhatib, Paul Aylin, Azeem Majeed et al.
Journal of Medical Internet Research
Blood Pressure and Hypertension Studies
article

Evaluation of a National Health Service Machine-Learning Model for Hypertension Case-Finding: Retrospective Cohort Study

Thomas Beaney, Ahmad Alkhatib, Paul Aylin, Azeem Majeed, Vesselin Novov, Thomas Woodcock, Gloria Ihenetu
article en

Abstract

Abstract Background Hypertension is a leading preventable cause of cardiovascular disease, yet a substantial proportion of adults remain undiagnosed, limiting opportunities for early intervention. A predictive model was commissioned by the North West London (NWL) Integrated Care Board to identify undiagnosed hypertension. The model was developed using health records from the Whole Systems Integrated Care (WSIC) database. Objective We aimed to independently evaluate the predictive performance of the model as it would be encountered in deployment, how performance varied by demographic characteristics, and practical utility. Methods To evaluate the predictive model, we conducted a retrospective cohort study of 1,802,920 individuals aged 16 years or older, registered with a general practice in NWL, and with no prior diagnosis of hypertension from May 2023 to May 2024. We assessed the model’s predictions against recorded hypertension status using medical diagnoses and blood pressure records. Logistic regression models were used to assess the sensitivity and specificity of the model’s predictions by sociodemographic groups. We also compared the model’s performance against a more interpretable regression approach. Results The model yielded an overall sensitivity of 62.7% (95% CI 62.5-62.8) and specificity of 60.7% (95% CI 60.5-60.8). Positive predictive value ranged from 31.5% (95% CI 31.2-31.8) to 42.9% (95% CI 42.5-43.2), and negative predictive value ranged from 77.6% (95% CI 77.2-77.9) to 84.9% (95% CI 84.7-85.2). Sensitivity was higher in older adults and Black patients; specificity was higher in younger adults, female patients, and White patients. Overall, sensitivity was higher for those living in areas of higher socioeconomic deprivation, while specificity was lower. These effects plateaued in the 2 least deprived quintiles of deprivation, which were comparable in both sensitivity and specificity. Predictions varied by age, with 96.2% (58,951/61,281) of those aged 70 to 79 predicted to have hypertension, whereas 0.08% of those aged 20 to 39 were predicted to have the condition. The model’s performance was comparable with a more interpretable logistic regression model. Conclusions Despite the model’s relatively good performance for those without hypertension, the positive predictive value was low, and a significant proportion of true cases remained undetected. Furthermore, there was considerable variation in performance associated with demographic characteristics, suggesting tailored approaches to case-finding may be beneficial in ensuring equity across demographic groups. Especially given the importance of understanding possible biases in predictive models, we recommend that, where there is no loss in performance, more parsimonious, transparent models be selected for prediction in health care settings. The findings of this evaluation can guide the practical application of the model, inform enhancements, direct targeted screening initiatives, and support cost-benefit analyses for broader implementation to improve hypertension management.

Journal of Medical Internet ResearchVol. 28
Openalex Percentile: Top 10%
Blood Pressure and Hypertension Studies
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.