ABSTRACT Aim To develop and validate machine learning models using routinely available clinical and laboratory data in adults with highly suspected acute appendicitis and to assess their potential to reduce negative appendectomies. Methods A retrospective study was conducted including adult patients who underwent appendectomy for suspected acute appendicitis between January 2020 and June 2024. Histopathology was the reference standard. Logistic regression, random forest, balanced random forest, and gradient boosting models were developed. Decision thresholds were calibrated to achieve predefined sensitivity targets. Performance was assessed using the area under the receiver operating characteristic curve (AUC) and standard diagnostic metrics. Stability was evaluated using bootstrap resampling and explainability by SHAP analysis. Results A total of 623 patients were included, of whom 66 had a negative appendectomy. For appendicitis detection, logistic regression demonstrated the best performance with a mean AUC of 0.765. At the highest sensitivity thresholds, specificity remained low but identified a small subset of patients at minimal risk of appendicitis. For complicated appendicitis prediction, random forest performed the best with a mean AUC of 0.785. SHAP analysis identified inflammatory laboratory markers as the most influential predictors. Bootstrap analysis confirmed model stability. Conclusion Machine learning models based on clinical and laboratory variables may provide adjunctive decision support in surgically selected adults with suspected acute appendicitis and may help identify a small subgroup in whom immediate surgery could be reconsidered at carefully selected sensitivity thresholds. Prospective external validation using a similar study design is required before clinical implementation.
Maleš et al. (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: