Abstract The technological advancements in mass spectrometry-based proteomics and metabolomics have enabled large-scale studies, in which hundreds to thousands of samples are analyzed in batches across multiple instruments over extended periods. Consequently, there are inevitable technical variations introduced as batch effects in the results. To address these issues, we developed a pipeline named “omics batch correct (OBC)” that integrates optimized data preprocessing steps, including normalization, handling of missing values, and batch correction, with a two-tier quality control (QC) system designed for both proteomic and metabolomic data. The first-tier QC incorporates methods such as principal component analysis, t-distributed stochastic neighbor embedding, uniform manifold approximation and projection, relative standard deviation analysis, Pearson correlation, and principal variance component analysis. The second-tier QC focuses on comparing differentially expressed molecules, particularly those with known regulatory roles, before and after batch correction. We validated the OBC through comprehensive cross-validation using clinical proteomic and metabolomic datasets, demonstrating its superior performance in mitigating batch effects while preserving biologically significant variations. The OBC pipeline is accessible via a user-friendly web interface at https://zhljude.shinyapps.io/OBC-app.
Zheng et al. (Sat,) studied this question.