Calibration of commercial AI chatbots for cancer symptom recognition: A pilot methodology study.
Abstract
e13698 Background: AI chatbots are increasingly used for health queries, yet their calibration for cancer recognition remains uncharacterized. How should AI respond to users with varying health literacy presenting equivalent clinical scenarios? Methods: We developed 114 calibrated queries across 10 cancer types, three pretest probability scenarios (high, low, diagnostic trap), and three user personas (naive patient, sophisticated patient, expert clinician) presenting equivalent clinical scenarios with varying linguistic framing. Each query was submitted to Claude, ChatGPT, and Gemini (n = 342 evaluable responses). Responses were scored using a keyword-matching algorithm. Results: Algorithm-derived Expected Calibration Alignment was 76.0% overall. High pretest probability scenarios showed 89.6% alignment; low pretest showed 38.9%, suggesting frequent over-escalation. Response thoroughness varied by persona: expert queries elicited responses ~2x longer than naive queries across all platforms. Expert responses included pre-test probabilities and workup protocols; naive responses provided general advice to "see a doctor." Conclusions: Current AI chatbots calibrate response thoroughness to query sophistication rather than clinical need. This raises alignment questions with equity implications: should AI proactively provide detailed guidance to users who don't know to ask for it? Data available via OSF.
Article Details
Journal Info
Journal of Clinical Oncology
Lippincott Williams & Wilkins
Authors (3)
Jennifer M. Hinkel
University of Oxford, Oxford, United Kingdom
Taylor Hirschberg
University of Oxford, Oxford, United Kingdom
Cory Kidd
Advient Advisors, Berkeley, CA