PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
March 7, 20260 citationsOpen Access

Tipping the Balance: Impact of Class Imbalance Correction on the Performance of Clinical Risk Prediction Models

View Full Paper
AAAmalie Koch AndersenHMHadi MehdizavarehAKArijit Khan

Key Points

  • The study investigates the effects of class-imbalance correction techniques on the performance of machine learning-based clinical risk prediction models.
  • Analyzed ten clinical datasets involving 605,842 patients
  • Evaluated multiple machine-learning model families, including linear and non-linear approaches
  • Applied three 1:1 class-imbalance correction strategies (SMOTE, RUS, ROS)
  • Assessed model performance using discrimination and calibration metrics on held-out data
  • Resampling had no positive impact on predictive performance across all datasets
  • Small and inconsistent changes in ROC-AUC observed, with some strategies showing negative effects
  • Models using imbalance correction had higher Brier scores, reflecting poorer probabilistic accuracy
  • Calibration performance degraded, with significant deviations in calibration intercept and slope

Abstract

Objective: ML-based clinical risk prediction models are increasingly used to support decision-making in healthcare. While class-imbalance correction techniques are commonly applied to improve model performance in settings with rare outcomes, their impact on probabilistic calibration remains insufficiently understood. This study evaluated the effect of widely used resampling strategies on both discrimination and calibration across real-world clinical prediction tasks. Methods: Ten clinical datasets spanning diverse medical domains and including 605,842 patients were analyzed. Multiple machine-learning model families, including linear models and several non-linear approaches, were evaluated. Models were trained on the original data and under three commonly used 1:1 class-imbalance correction strategies (SMOTE, RUS, ROS). Performance was assessed on held-out data using discrimination and calibration metrics. Results: Across all datasets and model families, resampling had no positive impact on predictive performance. Changes in the Receiver Operating Characteristic Area Under Curve (ROC-AUC) relative to models trained on the original data were small and inconsistent (ROS: -0.002, p0.05; SMOTE: -0.01, p<0.05), with no resampling strategy demonstrating a systematic improvement. In contrast, resampling in general degraded the calibration performance. Models trained using imbalance correction exhibited higher Brier scores (0.029 to 0.080, p<0.05), reflecting poorer probabilistic accuracy, and marked deviations in calibration intercept and slope, indicating systematic distortions of predicted risk despite preserved rank-based performance. Conclusion: In a diverse set of real-world clinical prediction tasks, commonly used class-imbalance correction techniques did not provide generalizable improvements in discrimination and were associated with degraded calibration.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Andersen et al. (2026) studied this question.

synapsesocial.com/papers/69abc0925af8044f7a4e93fdhttps://doi.org/10.48550/arxiv.2603.00208
Ask AI
Helpful
Bookmark
Share
View Full Paper