PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
April 30, 2026Journal of Clinical Medicine1 citationsOpen Access

Natural Language Processing of Unstructured Healthcare Data for Predicting Heart Failure in Individuals with Type 2 Diabetes

View Full Paper
JNJuan F. Navarro-GonzálezLIL Perez De IslaGMGloria Cánovas Molina

Key Points

  • To develop a predictive model for two-year heart failure risk in individuals with Type 2 Diabetes using unstructured healthcare data.
  • Retrospective multicenter study involving individuals with Type 2 Diabetes from eight hospitals in Spain.
  • Data extracted from electronic health records using clinical natural language processing.
  • Comparative analysis of logistic regression, random forest, and other algorithms for predictive modeling.
  • 14.3% of 588,756 individuals had prevalent heart failure.
  • iHF occurred in 13.6% of the training set and 11.4% in the validation set.
  • Logistic regression achieved the best AUC-ROC performance at 0.73 with 27 predictors.

Abstract

Background/Objectives: Type 2 diabetes mellitus (T2DM) is a multisystemic disease with overlapping metabolic, renal, and cardiovascular effects. Within the Diabetic@ project, which aims to characterize individuals with T2DM using real-world data extracted from electronic health records (EHRs), this substudy sought to develop a predictive model for two-year heart failure (HF) risk. Methods: Multicenter, retrospective study including T2DM individuals across eight Spanish hospitals (2013–2018). Data were extracted exclusively from EHRs’ unstructured free text using clinical natural language processing (cNLP) and mapped to SNOMED CT. At inclusion, individuals were categorized as having or not prevalent HF (pHF). Predictive modeling was performed in non-pHF to assess two-year risk of developing HF, termed incident HF (iHF). Logistic regression (LR), decision trees, random forest, and XGBoost were compared, selecting for accuracy and interpretability. Results: Of 588,756 individuals with T2DM, 84,197 (14.3%) had pHF. Among non-pHF, 353,371 (60%) were used for model development (90.7% training, 9.3% validation). iHF occurred in 13.6% of the training set and 11.4% of the validation set. Ischemic heart disease was present in 16.2% overall, 37.9% in pHF, and 12.6% in non-pHF. Glycosylated hemoglobin data was rarely reported (<15%). LR achieved the best performance (AUC-ROC 0.73) using 27 predictors. Reduced 12- and clinically refined 9-predictor models performed similarly, with the latter implemented in a web-based tool. Conclusions: Unstructured data from EHRs enabled development of a two-year HF risk model for individuals with T2DM, underscoring the potential of cNLP for risk stratification across the cardiovascular–renal–metabolic spectrum.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Navarro-González et al. (2026) studied this question.

synapsesocial.com/papers/69f2a42a8c0f03fd677632bahttps://doi.org/10.3390/jcm15093287
Ask AI
Helpful
Bookmark
Share
View Full Paper