Synthetic data generation from breast cancer patients to enable PSM on wide and multidimensional real-world populations.
Abstract
e12574 Background: In several surgical oncology domains randomized controlled trials are no longer feasible due to ethical and acceptability constraints. Evidence generation therefore relies on real-world data (RWD), which are intrinsically affected by strong selection bias. Propensityscore matching (PSM) improves comparability but, in highly multidimensional datasets, often causes substantial loss of sample size. Synthetic data generation may enable expansion of matched cohorts while preserving real-world structure. We tested this methodology in the historical setting of breast conservation and radiotherapy vs. mastectomy. Methods: SEER-based RWD were used to compare PSM performance between real and synthetic cohorts. A synthetic cohort was generatedusing a conditional generative adversarial network, preconditioned on the same covariates used for propensity score estimation, to preserve the full multidimensional joint distribution of clinical, demographic, and socioeconomic variables, including rare and extreme treatment patterns. Fidelity, correlation structure, utility, and privacy were validated using the Synthetic Validation Framework (SAFE). Propensity scores were estimated via logistic regression including 12 covariates. Stratified nearest-neighbor PSM was applied using distances computed in a whitened covariate space (Mahalanobis-equivalent). Matching quality was assessed using standardized mean differences (SMD; < 0.1 optimal, < 0.2 acceptable). Identical matching parameters were applied to real and synthetic datasets, followed by caliper sensitivity analyses. Results: When PSM was applied to RWD alone (22,405 patients), only 7.5% of MX-RT patients (601) were matched, despite excellent balance (mean SMD 0.02; all SMD < 0.2). The synthetic cohort (137,782 patients) substantially increased matching yield. Using identical parameters, 33.7% of MX-RT patients (17,383) were matched with preserved balance (mean SMD 0.037). After caliper optimization (c = 2.6), matching efficiency increased to 46.1% (23,753 patients), while maintaining acceptable balance (mean SMD 0.056; all SMD < 0.2). The treated-to-control ratio remained stable (~1:1.4). Overall, synthetic data enabled up to a six-fold increase in matched patients, allowing exploration of clinically relevant subgroups. Conclusions: Synthetic data generation, validated with SAFE, enables recovery of sample size lost to PSM in highly multidimensional RWD while preserving covariate balance and structural complexity. This approach does not generate new clinical evidence but enables adequately powered exploratory and subgroup analyses in observational settings where randomization is no longer feasible.
Article Details
Journal Info
Journal of Clinical Oncology
Lippincott Williams & Wilkins
Authors (9)
Giuseppe Catanuto
Humanitas University - Humanitas Istituto Clinico Catanese, Misterbianco, Catania, Italy
Saverio D'Amico
1IRCCS Humanitas Research Hospital, AI Center, Rozzano, Italy
Sara Baccino
IRCCS Humanitas Research Hospital, Rozzano, Italy
Alessandro Bruseghini
1IRCCS Humanitas Research Hospital, AI Center, Rozzano, Italy
Damiano Gentile
Humanitas University - IRCCS Humanitas Research Hospital, Rozzano, Milan, Italy
Federica Martorana
Francesco Caruso
Gaetano Castiglione
Humanitas Istituto Clinico Catanese, Misterbianco, Catania, Italy
Corrado Tinterri
Humanitas University - IRCCS Humanitas Research Hospital, Rozzano, Milan, Italy