PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
June 2, 2026BMC Endocrine Disorders0 citationsOpen Access

Electronic health record-derived machine learning model for hypoglycemia risk prediction in type 2 diabetes mellitus patients: development and validation

QRQian RanXQXia QiLLLi Liu

Key Points

  • The study aims to create and validate machine learning models for predicting hypoglycemia risk in patients with type 2 diabetes mellitus.
  • Cohort study design with data collected from electronic health records.
  • Data was randomly split into training (70%) and validation (30%) subsets.
  • Four machine learning algorithms (LR, XGBoost, RF, SVM) were used to predict hypoglycemia risk.
  • Hypoglycemia incidence in the study cohort was 22.0%.
  • In the validation cohort, XGBoost had the highest AUC of 0.78 compared to RF (0.75), SVM (0.72), and LR (0.76).
  • Key predictors identified were creatinine, triglycerides, albumin, HbA1c, C-peptide, AST, hemoglobin, and sulfonylurea use.

Abstract

BACKGROUND: Hypoglycemia is a serious complication of diabetes. Early recognition of hypoglycemia can improve clinical prognosis, however, traditional diagnostic tools are often limited. Machine learning offers a promising approach for predicting adverse outcomes in diabetic patients. OBJECTIVE: This study aims to develop and validate machine learning-based models to predict the risk of hypoglycemia in type 2 diabetes mellitus (T2DM) patients. METHODS: A cohort study design was employed. Clinical data were collected from the electronic health record system. The dataset was randomly partitioned into training and validation subsets using a 7:3 ratio. Four machine learning algorithms, logistic regression (LR), Extreme Gradient Boosting (XGBoost), random forest (RF), and support vector machine (SVM) were implemented to develop hypoglycemia risk prediction models. Predictive performance was assessed using sensitivity, specificity, accuracy, precision, F1 score, and the area under the receiver operating characteristic curve (AUC). RESULTS: 831 T2DM patients were included, the hypoglycemia incidence was 22.0%. In the training cohort, the AUC for the LR, XGBoost, SVM, and RF models were 0.82, 0.86, 0.84, and 0.80, and corresponding AUCs were 0.76, 0.78, 0.72, and 0.75 in the validation cohort. The XGBoost demonstrated the highest overall predictive performance. Feature importance analysis based on the XGBoost model identified creatinine, triglycerides, albumin, HbA1c, C-peptide, aspartate aminotransferase, hemoglobin, and sulfonylurea use as the most influential predictors of hypoglycemia risk. CONCLUSIONS: The XGBoost model exhibited superior predictive performance for achieving the higher AUC, F1 score, greater accuracy, sensitivity and specificity. This model enables effective identification of T2DM patients who may require intensified monitoring or targeted interventions to prevent hypoglycemic events. CLINICAL TRIAL NUMBER: Not applicable.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Ran et al. (2026) studied this question.

synapsesocial.com/papers/6a1e723f30b38c64201b58ebhttps://doi.org/10.1186/s12902-026-02340-9
Ask AI
Helpful
Bookmark
Share
View Full Paper