Molecular classification guides breast cancer treatment, but PAM50 and immunohistochemistry (IHC) remain costly and unavailable in many settings. Foundation models (FMs) combined with multiple instance learning (MIL) show promise for predicting molecular subtypes from haematoxylin-and-eosin-stained slides, yet most studies report only internal validation. This study evaluates FMs with MIL across cohorts and identifies factors associated with domain-induced performance degradation. We evaluate 13 FMs and 3 complementary MIL architectures for PAM50 subtyping and IHC biomarker prediction using cross-validation on TCGA-BRCA (Formula: see text) and external validation on CPTAC-BRCA (Formula: see text). Virchow v2 achieves the best overall performance but exhibits severe degradation upon external validation, consistent across all three MIL architectures especially for HER2-enriched and Normal-like PAM50 subtypes and HER2-positive IHC prediction. Four hypothesised domain shift factors are quantified through exploratory regression analysis to explain relative performance drop (RPD). Staining variability, feature space divergence and morphological separability reach significance in univariate analysis, whilst prevalence shift does not. Staining variability and feature space divergence as covariate-level factors jointly account for 80.0% of RPD variance in the most parsimonious multivariate model (Formula: see text, Formula: see text). Although based on a limited number of class-level observations and therefore exploratory in nature, these findings highlight the need for domain generalisation strategies targeting covariate shift, even when specialised FMs are used as feature encoders.
Fernandez-Romero et al. (2026) studied this question.