Machine learning and deep learning–based prediction of hypertension and analysis of its major risk factors in Bangladesh

BACKGROUND: Hypertension is a leading cause of cardiovascular morbidity and mortality in Bangladesh. This study examined its prevalence, risk factors, and predictive modeling using machine learning (ML) and deep learning (DL) approaches. METHOD: We analyzed cross-sectional data from the 2022 Bangladesh Demographic and Health Survey, which included 14,283 adults (≥18 years). Prevalence was estimated, chi-square tests assessed associations, and four ML models (weighted logistic regression, random forest, extreme gradient boosting, light gradient boosting machine) and two DL models (TabNet, and multi-layer perceptron) were applied to predict hypertension risk. Model performance was evaluated using accuracy, precision, recall, specificity, F1 score, and area under the receiver operating characteristics curve and precision-recall curve. RESULTS: Overall prevalence was 18.04% (95% CI: 17.2%-18.9%), higher among women (18.87%) than men (16.97%). The chi-square test suggests that hypertension was significantly associated with age, BMI, diabetes, wealth index, education, household size, and region (p < 0.05). Among the machine learning and deep learning models, weighted logistic regression (WLR) achieved the highest accuracy (0.817), precision (0.444), specificity (0.981), AUC-ROC (0.751), and AUC-PR (0.357). However, WLR exhibited low recall (0.070). In contrast, the random forest (RF) model achieved the highest recall (0.687) and F1-score (0.460) on the test data, indicating greater sensitivity in identifying individuals with hypertension. Additionally, age, BMI, sex, family size, and educational level were identified as the most important predictors among the variables included in the study. CONCLUSION: Hypertension is common in Bangladesh, with higher prevalence in women and significant association with socio-demographic determinants. Although WLR demonstrated the highest accuracy, precision, specificity, and AUC-PR, its low recall limits its utility for identifying individuals with hypertension. RF may be more suitable for public health applications because of its higher recall and F1-score; however, further external validation and assessment of its clinical utility are required before implementation.

Authors

Institutions

Publication Details

Journal
PLoS ONE
Published
2026-09-17
DOI
https://doi.org/10.1371/journal.pone.0358471
Primary Topic
Blood Pressure and Hypertension Studies
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Machine learning and deep learning–based prediction of hypertension and analysis of its major risk factors in Bangladesh

Samiul Islam, Md. Matiur Rahaman, M. M. Imran Molla, Mohammad Ali et al.
PLoS ONE
Blood Pressure and Hypertension Studies
article

Machine learning and deep learning–based prediction of hypertension and analysis of its major risk factors in Bangladesh

Samiul Islam, Md. Matiur Rahaman, M. M. Imran Molla, Mohammad Ali, Shawrab Chandra, Md. Ayub Ali
article en

Abstract

BACKGROUND: Hypertension is a leading cause of cardiovascular morbidity and mortality in Bangladesh. This study examined its prevalence, risk factors, and predictive modeling using machine learning (ML) and deep learning (DL) approaches. METHOD: We analyzed cross-sectional data from the 2022 Bangladesh Demographic and Health Survey, which included 14,283 adults (≥18 years). Prevalence was estimated, chi-square tests assessed associations, and four ML models (weighted logistic regression, random forest, extreme gradient boosting, light gradient boosting machine) and two DL models (TabNet, and multi-layer perceptron) were applied to predict hypertension risk. Model performance was evaluated using accuracy, precision, recall, specificity, F1 score, and area under the receiver operating characteristics curve and precision-recall curve. RESULTS: Overall prevalence was 18.04% (95% CI: 17.2%-18.9%), higher among women (18.87%) than men (16.97%). The chi-square test suggests that hypertension was significantly associated with age, BMI, diabetes, wealth index, education, household size, and region (p < 0.05). Among the machine learning and deep learning models, weighted logistic regression (WLR) achieved the highest accuracy (0.817), precision (0.444), specificity (0.981), AUC-ROC (0.751), and AUC-PR (0.357). However, WLR exhibited low recall (0.070). In contrast, the random forest (RF) model achieved the highest recall (0.687) and F1-score (0.460) on the test data, indicating greater sensitivity in identifying individuals with hypertension. Additionally, age, BMI, sex, family size, and educational level were identified as the most important predictors among the variables included in the study. CONCLUSION: Hypertension is common in Bangladesh, with higher prevalence in women and significant association with socio-demographic determinants. Although WLR demonstrated the highest accuracy, precision, specificity, and AUC-PR, its low recall limits its utility for identifying individuals with hypertension. RF may be more suitable for public health applications because of its higher recall and F1-score; however, further external validation and assessment of its clinical utility are required before implementation.

PLoS ONEVol. 21(9)
Khulna University (BD), Chhatrapati Shahu Ji Maharaj University (IN), University of Barishal (BD), University of Rajshahi (BD)
Good health and well-being
Openalex Percentile: Top 11%
Blood Pressure and Hypertension Studies
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.