Stepwise AI prompting for multiple myeloma imaging classification.
Abstract
e13707 Background: Electronic Health Records (EHRs) contain numerous variables, much of which are unstructured narrative data, posing challenges for systematic analysis. Large language models (LLMs) can help structure this data but are prone to hallucinations, limiting their reliability in clinical settings. This study evaluates ChatGPT’s ability to classify oncology imaging reports for the SLiM-CRAB criteria, a Multiple Myeloma diagnostic framework that relies on imaging findings like bone lesions. To improve accuracy and reduce errors, few-shot learning, and structured prompt engineering were used, which guided the process by using minimal labeled examples and breaking tasks into smaller steps. Methods: A stepwise prompting strategy was applied, dividing classification into (1) study type, (2) filtering unrelated records, and (3) identifying bone lesion presence. Depending on the study type, the presence of bone lesions fulfills either the “MRI” or “Bone Lesion” criterion in SLiM CRAB for Multiple Myeloma. Few-shot prompting (20–30 examples per step) improved accuracy and a temperature of 0 eliminated randomness from the response. Steps one and three utilized the OpenAI’s 4o model, while step two used the o1 model. For testing, a manually classified dataset of 216 unique imaging anonymized narratives was annotated by two domain experts; for medical decision-making, the standard AI performance benchmark was set at κ = 0.85. Cohen’s kappa was calculated for AI classifications, and McNemar’s Test was used to determine if discrepancies between AI and human classifications were statistically significant. Discrepancies were adjudicated through blinded secondary review by domain experts to determine the presence of human classification errors. Results: For step one, study type classification significantly outperformed human inter-rater reliability (κ = 0.98, p < 0.01); after adjudication, it reached 99% accuracy. When filtering unrelated records in step two (κ = 0.96), the concordance rate also outperformed the benchmark significantly (p < 0.01), reaching 98% accuracy after adjudication. For those that can be used for the SLiM CRAB Criteria (113), bone lesion detection in step three (κ = 0.89) outperformed the benchmark significantly (p < 0.05), after adjudication performance revealed 97% accuracy. The total accuracy rate for the three steps was 94%. Conclusions: A structured, multi-step prompting approach improves classification accuracy and reduces hallucinations by simplifying coding and improving the debugging of complex steps. ChatGPT-4o outperforms the medical decision-making standard AI performance benchmark annotation in study type classification, filtering of unrelated records, and final detection of bone lesions per the SLiM CRAB criteria. Further refinements, including rule-based validation and larger expert-curated databases for testing, are needed to improve context-sensitive classification.
Article Details
Journal Info
Journal of Clinical Oncology
Lippincott Williams & Wilkins
Authors (6)
Jay R Hydren
HealthTree Foundation, Lehi, UT
Alex C. Brown
HealthTree Foundation, Lehi, UT
Juan Pablo Capdevila
1HealthTree Foundation, South Jordan, United States
Jennifer M. Ahlstrom
HealthTree Foundation, Lehi, UT
Felipe Flores Quiroz
2HealthTree Foundation, South Jordan, United States
Jorge Arturo Hurtado Martinez
1HealthTree Foundation, South Jordan, United States