BACKGROUND: Differential gene selection is fundamental to transcriptomics; however, mainstream methods typically fit each gene independently via marginal analyses. Although multifactor designs can be accommodated, gene-gene interactions are not explicitly modeled, rendering results susceptible to systematic bias driven by coexpression patterns. Additionally, heterogeneous tools, inconsistent metrics, and fragmented workflows compromise the reproducibility and interpretability of downstream enrichment and visualization analyses. FINDINGS: We developed TransPro, an open-source integrated framework comprising two complementary packages, TransProPy and TransProR, for systematic benchmarking, bias correction, and reproducible visualization in differential gene selection. TransProPy (Python) combines multivariate AUC-based complementarity quantification with ensemble learning for interaction-aware gene selection, whereas TransProR (R) provides standardized differential analysis, pathway enrichment, and seven visualization workflows (circular dendrograms, chord diagrams, spiral plots, and interaction networks). Cross-dataset generalizability was assessed using 12 independent datasets spanning multiple cancer types and normal tissues with broad heterogeneity in data origin, platform, batch structure, and sample size. Under stringent thresholds, positively and negatively correlated genes maintained a near-equal proportion at the gene level, and activated and suppressed pathway proportions exhibited good concordance with gene-level correlation patterns; critically, TransProPy produced meaningful enrichment results under conditions where conventional methods failed. Quantitative comparisons of core enriched gene proportions further supported these differences (Kruskal-Wallis and pairwise Wilcoxon tests, p < 0.001). CONCLUSIONS: TransPro establishes a unified, reproducible framework that corrects method-specific biases while bridging computational discovery and biological interpretation. All code, workflows, documentation, and example data are openly accessible to support reproducibility and community reuse.
Yu et al. (Fri,) studied this question.