Key points are not available for this paper at this time.
BACKGROUND: Improved risk prediction of cancer-associated venous thromboembolism (VTE) remains an unmet clinical need. We aimed to apply machine learning for the prediction of cancer-associated VTE. PATIENTS AND METHODS: Data from the Vienna Cancer and Thrombosis Study (Vienna-CATS), a prospective cohort study that included patients with cancer from 2003-2019, were used. Patient characteristics and laboratory measurements (including various routine and experimental laboratory assays) were used to train and validate six classification models for VTE prediction. Monte Carlo cross-validation was conducted, with 80% of the samples randomly assigned to the training set and 20% to the test set. Preprocessing algorithms fitted on the training data were applied to the test data to avoid data leakage. Discrimination and input parameter influence were assessed. RESULTS: In total, 2193 patients (46.6% women) with a median age of 62 (interquartile range 52-68) years were included. The most common tumor types were lung (18.1%), brain (15.3%), and breast cancer (15.1%). Within 6 months and 2 years, 124 (cumulative incidence: 6.4%) and 186 (10.7%) VTE events occurred. The best performing model for VTE prediction was machine learning-enhanced Logistic Regression (LGR) with an area under the curve of 0.66 95% confidence interval (CI) 0.65-0.67 at 6 months and 0.62 (95% CI 0.61-0.62) at 2 years-comparable to established risk assessment models (RAMs). In LGR, the strongest contributors to VTE risk were homozygous factor V Leiden mutation, rectal, testicular, and pancreatic cancer, newly diagnosed cancer, and history of VTE. CONCLUSION: Machine learning models incorporating extensive clinical and biomarker panels did not outperform existing RAMs for cancer-associated VTE.
Hoberstorfer et al. (2026) studied this question.