Synthetic data generation increased the propensity score matching yield of breast cancer patients from 7.5% using real-world data alone to 46.1%, while maintaining acceptable covariate balance.
Observational (n=22,405)
Synthetic data generation can recover sample size lost to propensity score matching in highly multidimensional real-world datasets while preserving covariate balance.
Effect estimate: six-fold increase
Absolute Event Rate: 46.1% vs 7.5%
e12574 Background: In several surgical oncology domains randomized controlled trials are no longer feasible due to ethical and acceptability constraints. Evidence generation therefore relies on real-world data (RWD), which are intrinsically affected by strong selection bias. Propensityscore matching (PSM) improves comparability but, in highly multidimensional datasets, often causes substantial loss of sample size. Synthetic data generation may enable expansion of matched cohorts while preserving real-world structure. We tested this methodology in the historical setting of breast conservation and radiotherapy vs. mastectomy. Methods: SEER-based RWD were used to compare PSM performance between real and synthetic cohorts. A synthetic cohort was generatedusing a conditional generative adversarial network, preconditioned on the same covariates used for propensity score estimation, to preserve the full multidimensional joint distribution of clinical, demographic, and socioeconomic variables, including rare and extreme treatment patterns. Fidelity, correlation structure, utility, and privacy were validated using the Synthetic Validation Framework (SAFE). Propensity scores were estimated via logistic regression including 12 covariates. Stratified nearest-neighbor PSM was applied using distances computed in a whitened covariate space (Mahalanobis-equivalent). Matching quality was assessed using standardized mean differences (SMD; < 0.1 optimal, < 0.2 acceptable). Identical matching parameters were applied to real and synthetic datasets, followed by caliper sensitivity analyses. Results: When PSM was applied to RWD alone (22,405 patients), only 7.5% of MX-RT patients (601) were matched, despite excellent balance (mean SMD 0.02; all SMD < 0.2). The synthetic cohort (137,782 patients) substantially increased matching yield. Using identical parameters, 33.7% of MX-RT patients (17,383) were matched with preserved balance (mean SMD 0.037). After caliper optimization (c = 2.6), matching efficiency increased to 46.1% (23,753 patients), while maintaining acceptable balance (mean SMD 0.056; all SMD < 0.2). The treated-to-control ratio remained stable (~1:1.4). Overall, synthetic data enabled up to a six-fold increase in matched patients, allowing exploration of clinically relevant subgroups. Conclusions: Synthetic data generation, validated with SAFE, enables recovery of sample size lost to PSM in highly multidimensional RWD while preserving covariate balance and structural complexity. This approach does not generate new clinical evidence but enables adequately powered exploratory and subgroup analyses in observational settings where randomization is no longer feasible.
Catanuto et al. (Thu,) conducted a observational in Breast cancer (n=22,405). Synthetic data generation vs. Real-world data alone was evaluated on Matching yield of MX-RT patients (six-fold increase). Synthetic data generation increased the propensity score matching yield of breast cancer patients from 7.5% using real-world data alone to 46.1%, while maintaining acceptable covariate balance.