Abstract Background Hospital‐acquired venous thromboembolism (HA‐VTE) is a significant cause of morbidity and mortality among hospitalized adults. Accurate prediction of HA‐VTE is crucial for timely intervention and prevention. While logistic regression is widely used for the development of clinical prediction models, there is ongoing interest in the potential for machine learning methods to enhance prediction performance. Objectives This study aimed to identify the most practical method for the prediction of HA‐VTE based on electronic health record (EHR) data. Methods We evaluated and compared the performance of prognostic models developed using logistic regression, random forest, extreme gradient boosting (XGBoost), deep neural networks, and an ensemble method to predict HA‐VTE using EHR data from a large academic medical center. Models were evaluated in a temporal external validation cohort based on discrimination and calibration metrics, overall and among key patient subgroups. Results All models demonstrated similarly high discrimination, with C statistics ranging from 0.886 to 0.900. Logistic regression, random forest, and XGBoost showed excellent calibration, whereas deep neural networks exhibited poorer calibration. Additionally, prediction accuracy at a probability cut‐off of 0.02 and correlations between models indicated that logistic regression performed comparably well, if not better, compared with machine learning methods. Subgroup analyses further confirmed that the logistic regression model demonstrated strong discrimination and calibration among patient subpopulations. Conclusions Our findings suggest that logistic regression is a highly effective tool for EHR‐based prediction of HA‐VTE.
Ko et al. (2026) studied this question.