Personality-driven roles of large language models for axillary surgery recommendation.

C Claire Dvorak (Orlando Health, Orlando, FL) T Tomas Dvorak J Jeffrey R. Smith (Orlando Health, Orlando, FL) D Danielle Henry

Abstract

e13703 Background: Large language models (LLMs) are increasingly evaluated for clinical decision support. Inclinical practice, surgical oncologists vary in decision style, tolerance of uncertainty, and willingness tocommit with incomplete information. We hypothesized that LLMs may exhibit analogous“personality” profiles with different strengths and weaknesses. Methods: Oncologic history narratives from 100 breast cancer cases with adjudicated axillary surgeries (sentinellymph node biopsy [SLNB], axillary lymph node dissection [ALND], or no axillary surgery [NONE]) wereprovided verbatim to four LLMs (Gemini-3, GPT-5.2, Claude-4.5, Grok-4). Axillary surgery details wereremoved from the input. Models were prompted to recommend upcoming SLNB, ALND, or NONE, withabstention permitted. Two epistemic regimes were tested: assumption-permissive (Arm A; inference ofunspecified negatives), and assumption-conservative (Arm B, prohibiting assumptions beyond text).Coverage cases with definitive recommendation, conditional concordance the agreement with observedsurgery among non-abstaining cases. Multi-model ensembles (3-of-4, 4-of-4) were evaluated. All datawere de-identified and analyzed under IRB. Results: LLMs demonstrated clinically distinct decision-support profiles. Gemini exhibited a “permissive generalist”profile, maintaining high coverage across regimes from Arm A to Arm B (0.96→0.92) with stableconcordance (0.62→0.64), consistent with high-throughput triage behavior. Grok demonstrated a“cautious specialist” profile, with reduced coverage (0.70→0.44) but highest concordance whencommitting (0.71→0.79), consistent with selective validation. Claude showed an “over-cautious” profile,reducing coverage (0.89→0.59) without improvement in concordance (0.40→0.39), while GPT-5.2behaved as a “defensive clinician,” with the largest coverage decline (0.82→0.23) and only modestconcordance gains (0.27→0.43). Most common miss was recommendation of NONE (34%) or ABSTAIN(31%) when SLNB was performed.Ensemble strategies did not improve performance. Both 3-of-4 and 4-of-4 agreement substantiallyreduced coverage (≤0.55 in Arm A; ≤0.21 in Arm B) without exceeding the concordance of the bestindividual models. Conclusions: LLMs exhibit clinician-like personality profiles that support role-based deployment rather thaninterchangeable use. Permissive models may be suited for high-throughput triage (pre-clinic review),cautious models for high-confidence validation (decision second check), and abstention-prone models fordocumentation gaps (escalation for human review). Although observed concordance is insufficient forautonomous clinical decision-making, these patterns provide guidance for early-stage delegation ofsubtasks, in which LLMs function as complementary teammates to surgical oncologists across themanagement spectrum.

Article Details

Volume / Issue Vol. 44, Issue 16_suppl
Published June 01, 2026
ISSN 0732-183X
Publisher Lippincott Williams & Wilkins

Journal Info

Journal of Clinical Oncology

Lippincott Williams & Wilkins

ISSN: 0732-183X Health Sciences

Authors (4)

C

Claire Dvorak

Orlando Health, Orlando, FL

T

Tomas Dvorak

J

Jeffrey R. Smith

Orlando Health, Orlando, FL

D

Danielle Henry