Background: The Banff 2022 classification endorses intragraft gene-expression profiling using the Banff Human Organ Transplant (B-HOT) consensus gene panel for rejection diagnosis. However, lack of standardized analytical pipelines, including data normalization, limits clinical implementation, with its impact on diagnostic performance yet to be determined. Methods: We evaluated ten normalization methods in 868 kidney allograft biopsies from nine European and North American centers, all Banff-graded and B-HOT profiled on nCounter, comprising a derivation ( n =441), internal ( n =186) and external ( n =241) validation cohorts. Each method was assessed through its downstream impact on: (i) gene count stability, (ii) differential expression and cross-platform concordance with RNA-seq data, and (iii) discrimination and calibration of predictive models for antibody- (AMR) and T cell-mediated rejection (TCMR). Results: Most methods improved count stability and showed high concordance with RNA-seq for overall gene expression. They also produced robust differential expression signatures consistent with those detected by RNA-seq, except for RUVSeq and RCRNorm , which identified fewer differentially expressed genes and showed lower concordance. In the overall validation cohort ( n =427), diagnostic performance was consistently high across nSolver -based approaches, nanostringr , NanoStringDiff , MetaNorm , and RCRNormFast (AMR AUROC 0.88–0.91; AUPRC 0.86–0.89; TCMR AUROC 0.90–0.92; AUPRC 0.78–0.83). Performance declined with RCRNorm (AMR AUROC/AUPRC 0.55/0.41; TCMR 0.53/0.18) and, for TCMR, with RUVSeq (AUROC 0.84–0.85; AUPRC 0.64–0.65). Calibration was satisfactory for most methods, except for RCRNorm and for TCMR models after RUVSeq . Conclusions: Normalization choice significantly impacted gene expression profiles and diagnostic classifier performance. Most methods, including nSolver-based pipelines, achieved robust discrimination for both AMR and TCMR. Complex methods, including RCRNorm, and RUVSeq for TCMR, reduced performance, with simpler approaches consistently outperforming them for B-HOT-based molecular diagnostics.
Piedrafita et al. (2026) studied this question.