Evaluation of Gemini Pro for primary site prediction in a real-world blinded cohort of 465 patients.
Abstract
e15002 Background: Artificial intelligence (AI) integration in oncology has transitioned from experimental use to routine clinical practice. Large Language Models (LLMs), such as Gemini Pro, are increasingly used for real-time decision support and molecular interpretation. This study evaluates Gemini Pro’s ability to identify the primary tumor site—defined by histopathology—using only genomic alterations, age, and sex. Methods: We performed a blinded, retrospective analysis of 465 evaluable Next-Generation Sequencing (NGS) cases, including 356 tissue biopsies and 109 liquid biopsies, generated via comprehensive hybrid-capture sequencing. Gemini Pro was queried using a standardized zero-shot prompt to provide three ranked primary tumor site predictions with confidence levels (Very High, High, Moderate, Low) and molecular justification based on lineage-specific genomic features. No supervised training or model fine-tuning was performed. Performance was assessed using Top-1 and Top-3 accuracy. Pearson’s Chi-square tests compared subgroups, and multivariate logistic regression calculated adjusted odds ratios (aOR) with 95% confidence intervals (CI). Results: Gemini Pro achieved a Top-1 accuracy of 50.1% (233/465; 95% CI: 45.6–54.6%) and a Top-3 accuracy of 68.4% (318/465; 95% CI: 64.2–72.6%). Model-assigned confidence was the strongest predictor of correctness: “Very High” predictions achieved 81.6% accuracy (102/125) with an aOR of 9.46 (95% CI: 5.82–15.38; p < .001) versus Moderate/Low confidence. Performance was strongest in common tumor lineages, including Breast (74.6%, 53/71) and Colorectal (72.8%, 43/59), where genomic anchors such as GATA3 and APC are prevalent. Accuracy was lower in rare or genomically complex tumors, particularly sarcomas (< 20%), which frequently exhibited non-specific copy number alterations. No significant difference was observed between tissue (54.6%) and liquid biopsies (42.9%; p = .174). Discordant cases commonly reflected a “neuroendocrine trap,” with TP53/RB1-deficient tumors misclassified as small cell lung cancer regardless of origin. Conclusions: In this blinded evaluation, Gemini Pro demonstrated moderate accuracy as a genomic-based tumor-origin classifier, supporting its role as a diagnostic adjunct rather than a replacement for histopathology. High-confidence predictions increased reliability and may guide targeted immunohistochemical confirmation within established diagnostic workflows. Category Variable Top-1 Accuracy (%) Adjusted Odds Ratio (OR) p-value Global Performance Full Study Cohort 50.1% -- -- Top-3 Prediction Rate 68.4% -- -- Model Confidence Very High 81.6% 9.46 < 0.001 High 63.0% 2.60 < 0.001 Moderate / Low 27.8% Reference -- High-Accuracy Lineages Breast 74.6% 7.92 < 0.001 Colorectal 72.8% 4.89 < 0.001 Lung 58.8% 3.06 < 0.001 Specimen Type Tissue Biopsy 54.6% 1.38 0.174 (NS) Liquid Biopsy (ctDNA) 42.9% Reference --
Article Details
Journal Info
Journal of Clinical Oncology
Lippincott Williams & Wilkins
Authors (13)
Bayan Altalla
King Hussein Cancer Center, Amman, Jordan
Marwan S. Al-Akasheh
Private Sector, Amman, Jordan
Kamal Hosni Al-rabi
King Hussein Cancer Center, Amman, Jordan
Mohammed Al-Jaghbeer
King Hussein Cancer Center, Amman, Jordan
Omar shafeeq Al-Rawi
Istishari Hospital, Amman, Jordan
Yazan Hamadneh
Jordan University Hospital, Amman, Jordan
Nour Maher Mustafa
Jordan University Hospital, Amman, Jordan
Anas Mohammad Zayed
King Hussein Cancer Center, Amman, Jordan
Tamer Moh'd Waleed Al-Batsh
King Hussein Cancer Center, Amman, Jordan
Mohammad Ma'koseh
Department of Medical Oncology, King Hussein Cancer Center, Amman, ON, Jordan
Akram Al-Ibraheem
Osama El Khatib
King Hussein Medical Center, Amman, Jordan
Nour A. Obeidat
King Hussein Cancer Center, Amman, Jordan