Introduction: Unnecessary laboratory testing can harm pediatric patients. Models predicting expected ranges of complete blood count (CBC) components may inform test necessity. It is commonly assumed that model generalizability improves with more sites in the training data. We tested the hypothesis that single-site CBC models would generalize as well as multi-site models. Methods: We included ICU encounters with > 1 CBC test from 5 ICUs (2012–2018) in the PICU Data Collaborative. We derived 69 features from demographics and lab history. Four Extra Trees Quantile Regression models were trained – three single-site (A, B, C) and one composite (A + B + C) – to predict ranges for hemoglobin (HGB), platelet count (PLT), and white blood cell count (WBC) right before all CBC tests except for the first. Models were evaluated on two hold-out sites (D, E) using the Winkler score at 80% target coverage. The score penalizes both wide intervals and failure to capture the truth, with lower scores indicating better performance. Paired Winkler differences between single-site and composite models were described using the Hodges-Lehmann median shift estimates and tested for significance using the Wilcoxon signed-rank tests. Results: We analyzed 220,604 CBC tests from 51,807 ICU encounters, with 80.6% used for development (Site A: 39.4%, B: 17.0%, C: 24.2%) and 19.4% (D: 12.3%, E: 7.1%) for evaluation. All models achieved approximately 80% coverage for all components (77.4 - 81.0%). The site A model performed similarly to the composite model for WBC (median shift 95% CI: +0.01 -0.01,+0.03; p=0.46), while the other models performed worse (B: +0.35 +0.33, +0.38; C: +0.43 +0.41, +0.46; both p< 0.001). Similar findings were observed for PLT (A: -0.10 -0.40, +0.20, p=0.55; B: +3.75 +3.35, +4.10, p< 0.001; C: +6.90 +6.55, +7.20, p< 0.001). For HGB, the site B model outperformed the composite model (-0.09 -0.10, -0.07, p< 0.001), but the other models performed worse (A: +0.01 0.00, +0.03; C: +0.16 +0.15, +0.17; both p< 0.001). Conclusions: Multi-site CBC models did not consistently outperform single-site models. Mixed results suggest that training with more sites does not always improve model generalizability. Future work will explore site characteristics contributing to differences in generalizability.
Huang et al. (Sun,) studied this question.