Personality-driven roles of large language models for axillary surgery recommendation.
Abstract
e13703 Background: Large language models (LLMs) are increasingly evaluated for clinical decision support. Inclinical practice, surgical oncologists vary in decision style, tolerance of uncertainty, and willingness tocommit with incomplete information. We hypothesized that LLMs may exhibit analogous“personality” profiles with different strengths and weaknesses. Methods: Oncologic history narratives from 100 breast cancer cases with adjudicated axillary surgeries (sentinellymph node biopsy [SLNB], axillary lymph node dissection [ALND], or no axillary surgery [NONE]) wereprovided verbatim to four LLMs (Gemini-3, GPT-5.2, Claude-4.5, Grok-4). Axillary surgery details wereremoved from the input. Models were prompted to recommend upcoming SLNB, ALND, or NONE, withabstention permitted. Two epistemic regimes were tested: assumption-permissive (Arm A; inference ofunspecified negatives), and assumption-conservative (Arm B, prohibiting assumptions beyond text).Coverage cases with definitive recommendation, conditional concordance the agreement with observedsurgery among non-abstaining cases. Multi-model ensembles (3-of-4, 4-of-4) were evaluated. All datawere de-identified and analyzed under IRB. Results: LLMs demonstrated clinically distinct decision-support profiles. Gemini exhibited a “permissive generalist”profile, maintaining high coverage across regimes from Arm A to Arm B (0.96→0.92) with stableconcordance (0.62→0.64), consistent with high-throughput triage behavior. Grok demonstrated a“cautious specialist” profile, with reduced coverage (0.70→0.44) but highest concordance whencommitting (0.71→0.79), consistent with selective validation. Claude showed an “over-cautious” profile,reducing coverage (0.89→0.59) without improvement in concordance (0.40→0.39), while GPT-5.2behaved as a “defensive clinician,” with the largest coverage decline (0.82→0.23) and only modestconcordance gains (0.27→0.43). Most common miss was recommendation of NONE (34%) or ABSTAIN(31%) when SLNB was performed.Ensemble strategies did not improve performance. Both 3-of-4 and 4-of-4 agreement substantiallyreduced coverage (≤0.55 in Arm A; ≤0.21 in Arm B) without exceeding the concordance of the bestindividual models. Conclusions: LLMs exhibit clinician-like personality profiles that support role-based deployment rather thaninterchangeable use. Permissive models may be suited for high-throughput triage (pre-clinic review),cautious models for high-confidence validation (decision second check), and abstention-prone models fordocumentation gaps (escalation for human review). Although observed concordance is insufficient forautonomous clinical decision-making, these patterns provide guidance for early-stage delegation ofsubtasks, in which LLMs function as complementary teammates to surgical oncologists across themanagement spectrum.
Article Details
Journal Info
Journal of Clinical Oncology
Lippincott Williams & Wilkins
Authors (4)
Claire Dvorak
Orlando Health, Orlando, FL
Tomas Dvorak
Jeffrey R. Smith
Orlando Health, Orlando, FL
Danielle Henry