Validating a semantic similarity approach for automated data extraction in phase II oncology trials.

E Evelyn Herrmann (Radio-Onkologie Zentrum Biel Seeland Berner Jura, Biel/Bienne, Switzerland) L Lukas Urs Von Rohr (Hirslanden Klinik Linde, Biel/Bienne, Switzerland) F Fizi Kotta Parambil (UXP Enterprise Solutions, Thiruvananthapuram, Kerala, India) S Sarath M Joy (UXP Enterprise Solutions, Thiruvananthapuram, Kerala, India) Y Yojena Chittazhathu Kurian Kuruvilla (Hospital of Thun, Thun, Switzerland)

Abstract

e13666 Background: Manual data extraction from phase II oncology publications is error-prone, especially when text is paraphrased. Traditional exact-match metrics underestimate correctness. We developed a pipeline using GPT-4 and BERT to capture semantic meaning, aiming to improve large-scale data extraction. Methods: We analyzed 26 phase II oncology publications, each mapped to > 200 predefined fields (e.g., study ID, drug names, mechanisms of action, outcomes). Domain experts established ground-truth (“base file”) labels. PDFs were parsed via Azure Document Intelligence Studio (custom model), chunked by LangChain, and stored as text embeddings in a vector database. For each data field, a specialized prompt retrieved relevant text chunks, and GPT-4 generated the final output. Performance was compared using: Classification/Exact Match—A field was correct only if its extracted text closely matched the reference. Semantic Similarity—BERT-based cosine scores (0–1) for textual fields; exact or TF-IDF checks for numeric data. GPT-4 also provided token-level log probabilities (“confidence scores”), though no formal threshold was set. We deemed a similarity ≥0.80 as correct. Results: Across 26 publications, we evaluated 5876 fields (1846 textual; 4030 numeric). Using strict classification, 40% of textual fields matched. By contrast, the BERT-based approach (≥0.80 threshold) yielded 88% correct matches (±0.05), with sponsor and drug-name fields reaching 95% accuracy. Numeric fields achieved 90% exact match, and TF-IDF-based similarity averaged 0.88 (±0.04). Preliminary error analysis indicated that synonyms (e.g., for mechanisms of action) caused most discrepancies. Table 1 illustrates how some outputs fail strict matching yet exhibit high semantic similarity. Conclusions: Emphasizing meaning over phrasing, a GPT-4/BERT pipeline outperforms strict classification for extracting textual fields from oncology trials. Future work will incorporate expert scoring to validate alignment with human judgment. Sample comparison: Traditional vs. semantic metrics. Field Base File Value Extracted Output Accuracy Score Semantic Similarity Sponsor(s) BiPar Sciences, privately held US biopharmaceutical… BiPar Sciences, Sanofi-Aventis 0 0.7273 Drug_Name Iniparib Iniparib 1 1 Molecule_Type Irreversible PARP inhibitor PARP inhibitor, non-selective, irreversible 0 0.8135 Experimental_Arm Chemotherapy; Targeted Treatment Chemotherapy and Targeted Treatment 0 0.9137 Masking Open Label open-label 0 0.8918 Study_Type Interventional Interventional 1 1 Described_in_Paper 1 = yes Yes 0 0.7064 Statistical_Test_Used Kaplan–Meier; log-rank; Pearson chi-square test - Descriptive - Pearson chi-square - Kaplan-Meier… 0 0.907 Study_Power 80 80 1 1 Accuracy score = 1 if exact or near-verbatim match, else 0. Semantic similarity score (0–1) ≥0.80 is considered correct.

Article Details

Volume / Issue Vol. 43, Issue 16_suppl
Published June 01, 2025
ISSN 0732-183X
Publisher Lippincott Williams & Wilkins

Journal Info

Journal of Clinical Oncology

Lippincott Williams & Wilkins

ISSN: 0732-183X Health Sciences

Authors (5)

E

Evelyn Herrmann

Radio-Onkologie Zentrum Biel Seeland Berner Jura, Biel/Bienne, Switzerland

L

Lukas Urs Von Rohr

Hirslanden Klinik Linde, Biel/Bienne, Switzerland

F

Fizi Kotta Parambil

UXP Enterprise Solutions, Thiruvananthapuram, Kerala, India

S

Sarath M Joy

UXP Enterprise Solutions, Thiruvananthapuram, Kerala, India

Y

Yojena Chittazhathu Kurian Kuruvilla

Hospital of Thun, Thun, Switzerland