ABSTRACT Background With high colorectal cancer (CRC) incidence, accurate early differentiation of precancerous polyps is critical for prognosis; while the gold‐standard colonoscopy‐biopsy is limited by invasiveness, cost and poor scalability, and routine blood tests lack efficiency with simple indicator combinations, machine learning's feature‐mining capacity offers a solution. This study aimed to develop and evaluate a machine learning system for differentiating patients with CRC and colorectal polyps using routine blood indices. Methods A retrospective analysis was conducted on the clinical data of 284 patients with CRC and 79 patients with colorectal polyps who were diagnosed at the Chinese PLA General Hospital from October 2021 to February 2024. The extreme gradient boosting (XGBoost) algorithm was used to establish a machine learning model using demographic characteristics and routine blood indices. The Shapley additive explanation method was used to evaluate feature importance. Results The constructed XGBoost model achieved high levels in differentiating CRC and colorectal polyps, with a precision of 0.906, a recall of 0.817, an accuracy of 0.791, and an area under the receiver operating characteristic curve of 0.869. The Shapley additive explanation showed that the top five important features were fibrinogen, carcinoembryonic antigen, plasma thrombin time, ferritin, and D‐dimer. Conclusion The XGBoost machine learning model based on blood indices has certain application value in the differential diagnosis of CRC and colorectal polyps, providing a new efficient tool for auxiliary diagnosis to assist clinical decision‐making.
Yan et al. (2026) studied this question.