Logistic regression outperformed other machine learning models in predicting hypertension (ROC-AUC 0.729; 95% CI 0.677-0.779) and undiagnosed hypertension (ROC-AUC 0.596; 95% CI 0.537-0.654).
Cross-Sectional (n=1,603)
Logistic regression was the best-performing machine learning model for identifying risk factors of hypertension and undiagnosed hypertension in rural Bangladesh.
Effect estimate: ROC-AUC 0.729 (95% CI 0.677-0.779)
Abstract Background Hypertension is a major cause of death and disability, and undiagnosed cases are particularly dangerous as they can cause severe damage without timely treatment. The aim of the study was to identify risk factors for hypertension and undiagnosed hypertension in rural areas of Bangladesh using advanced Machine Learning (ML) algorithms. Methods This study involved 1,603 respondents, selected through a cross-sectional survey using a multistage cluster random sampling technique. Four ML algorithms, including Gradient Booster (GB), Logistic Regression (LR), Random Forest (RF) and Support Vector Machine (SVM), were used in this study. Risk factors for hypertension and undiagnosed hypertension were identified using the best-performing ML model, selected based on metrics such as accuracy, sensitivity, specificity, precision, F1 score, receiver operating characteristics-area under the curve (ROC-AUC), and calibration plot. Results The prevalence of hypertension was 15.5%, slightly higher than the 15.4% for undiagnosed hypertension. In predicting the risk of both hypertension and undiagnosed hypertension, the LR model outperformed other ML models across most evaluation metrics. For hypertension, it achieved higher performance in terms of precision (0.580), F1 score (0.550), ROC-AUC (0.729; 95% CI: 0.677–0.779), and calibration. Similarly, for undiagnosed hypertension, the LR model showed better precision (0.580), ROC-AUC (0.596; 95% CI: 0.537–0.654), and calibration compared to other models. The risk factors for hypertension and undiagnosed hypertension differed notably. Key risk factors for undiagnosed hypertension included being overweight or obese, the absence of chronic diseases or cardiovascular disease (CVD), being male, non-use of tobacco, older age (above 50 years), being currently married, non-smoking status, having diabetes, and having no formal education. Conclusion The findings emphasize the urgent need for enhanced national and regional public health initiatives to improve the detection and awareness of hypertension in rural Bangladesh. Further research is important to validate the findings.
Bornee et al. (2026) conducted a cross-sectional in Hypertension and undiagnosed hypertension (n=1,603). Machine Learning algorithms was evaluated on Prediction of hypertension (ROC-AUC 0.729, 95% CI 0.677-0.779). Logistic regression outperformed other machine learning models in predicting hypertension (ROC-AUC 0.729; 95% CI 0.677-0.779) and undiagnosed hypertension (ROC-AUC 0.596; 95% CI 0.537-0.654).