PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
May 7, 2026BMC Medical Informatics and Decision Making0 citationsOpen Access

Interpretable machine learning for postoperative nausea and vomiting prediction in elderly orthopedic patients: a comparative study

LLLi‐Heng LiHGHao GuoHWHao Wang

Key Points

  • This study aims to create a machine learning model to predict postoperative nausea and vomiting (PONV) in elderly orthopedic patients.
  • Included 1216 elderly patients undergoing elective hip or knee surgery.
  • Data was split into training, validation, and independent test sets before any imputation.
  • Developed a StackNet meta-model using hyperparameter optimization and assessed clinical utility through Brier scores and SHAP for interpretability.
  • PONV incidence was 33% overall.
  • The StackNet model achieved an AUC of 0.9338 and significantly outperformed Logistic Regression (AUC = 0.7564, p < 0.001).
  • On the independent test set, StackNet accuracy was 0.7860 with sensitivity of 0.9250 and specificity of 0.7178.

Abstract

BACKGROUND: Postoperative nausea and vomiting (PONV) prolongs hospitalization and reduces patient satisfaction. Identifying high-risk elderly patients requires accurate absolute risk assessments, yet existing tools often lack probability calibration and transparency. METHODS: We included 1216 elderly patients undergoing elective hip or knee surgery. To strictly prevent data leakage, the dataset was partitioned into training, validation, and independent test sets in a 7:1:2 ratio prior to any imputation or feature selection. Following the systematic hyperparameter optimization of 12 distinct machine learning algorithms, a StackNet meta-model was developed by fusing optimal base-learner probabilities with raw clinical features. Clinical utility was evaluated via Brier scores and Decision Curve Analysis (DCA), alongside SHapley Additive exPlanations (SHAP) interpretability. RESULTS: Overall PONV incidence was 33%. The StackNet model achieved an AUC of 0.9338, significantly outperforming the conventional Logistic Regression baseline (AUC = 0.7564, p < 0.001) with superior calibration (Brier score = 0.102). On the independent test set, the StackNet model achieved an accuracy of 0.7860, sensitivity of 0.9250, specificity of 0.7178, and AUC of 0.9338, while the Logistic Regression baseline achieved an accuracy of 0.6584, sensitivity of 0.6750, specificity of 0.6503, and AUC of 0.7564. SHAP analysis identified preoperative frailty status and baseline hemoglobin levels as primary risk drivers. CONCLUSION: The StackNet framework offers highly calibrated absolute risk estimates for PONV in elderly orthopedic patients. Combined with SHAP transparency, it provides a clinically actionable tool to facilitate personalized antiemetic prophylaxis while avoiding unnecessary medical interventions due to overestimated risks.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Li et al. (2026) studied this question.

synapsesocial.com/papers/69fc2ba98b49bacb8b347a97https://doi.org/10.1186/s12911-026-03527-9
Ask AI
Helpful
Bookmark
Share
View Full Paper