Predicting progression-free and overall survival from response endpoints in oncology clinical trials: An embedding-based machine learning approach.
Abstract
e13666 Background: Objective response rate (ORR) and duration of response (DoR) provide important early efficacy signals in oncology trials. Their ability to predict subsequent progression-free survival (PFS) and overall survival (OS) varies widely by tumor type and treatment mechanism. Prior endpoint surrogacy analyses have focused on tumor type and treatment-specific statistical correlations, limiting generalizability across heterogeneous settings. We developed a machine learning (ML) framework incorporating ORR, DoR, and LLM-based clinical and treatment embeddings to predict median PFS (mPFS) and median OS (mOS) in a comprehensive cross-tumor clinical trials dataset. Methods: Data were extracted from the LARVOL CLIN, an outcomes database (2004–2023), comprising 1,223 Phase I–III trials (2,086 experimental and control arms) reporting both ORR and DoR. Median PFS and OS were available for 908 and 690 trials, respectively. To account for cross-trial heterogeneity, key population descriptors (tumor type, stage, treatment setting, prior lines of therapy, biomarker status) were embedded separately to preserve granularity, while treatment details (drugs, modalities, and dose/schedule) were embedded into a single representation. Embeddings from Sentence Transformers (all-mpnet-base-v2 embeddings) were used and reduced to 30 principal components (PCs). Two gradient boosting algorithms (XGBoost [XGB], LightGBM) were evaluated to predict log-transformed mPFS and mOS using log (ORR), log (DoR), arm size, population embeddings and treatment embeddings. To prevent intra-trial data leakage, trial-grouped five-fold cross-validation (CV) was used. Results: The best-performing model, XGBoost, achieved higher predictive accuracy for mPFS than mOS (CV R²: 0.72 vs 0.53 on log scale). SHAP importance analysis revealed that mPFS predictions were primarily driven by ORR and DoR, followed by tumor type and stage. Conversely, mOS predictions were more heavily influenced by tumor type (exceeding ORR and DoR contributions), stage, and treatment description. Limitations include the modest dataset size and uncertain generalizability to rare tumor subtypes. Conclusions: Integrating early response endpoints with trial context features enabled prediction of survival outcomes across heterogeneous oncology trials, with higher accuracy for mPFS than mOS. The stronger influence of tumor type and treatment context on mOS prediction suggests that early response endpoints alone may be insufficient surrogates for OS. This framework may support early trial prioritization and survival endpoint estimation when mature OS data are unavailable. XGB performance (log scale). Endpoint R² RMSE MAE mPFS (n=908) 0.72 0.30 0.20 mOS (n=690) 0.53 0.32 0.22 R2 = Coefficient of Determination, RMSE = Root Mean Squared Error, MAE = Mean Absolute Error.
Article Details
Journal Info
Journal of Clinical Oncology
Lippincott Williams & Wilkins
Authors (4)
Ankit Kalucha
The Larvol Group, LLC, San Francisco, CA
Judith Pérez Granado
The Larvol Group, LLC, San Francisco, CA
Mark Gramling
The Larvol Group, LLC, San Francisco, CA
Bruno Larvol
The Larvol Group, LLC, San Francisco, CA