Uncertainty-aware AI triage for LUAD histologic subtyping: Flagging cases for expert review.
Abstract
e20018 Background: AI-assisted histological subtyping of lung adenocarcinoma (LUAD) shows promise but imperfect accuracy in the setting of histologic heterogeneity raises concerns about clinical deployment. Current AI systems provide predictions without indicating confidence, leaving clinicians unable to identify cases requiring expert review. Misclassification of high-risk patterns (e.g., solid and micropapillary) can affect risk stratification. We developed an uncertainty-aware triage framework that identifies difficult cases for pathologist review while allowing confident AI predictions to proceed, optimizing workflow efficiency. Methods: We analyzed 143 resected LUAD whole slide images with five predominant subtypes: acinar (n = 60, 42%), solid (n = 55, 38%), lepidic (n = 16, 11%), micropapillary (n = 10, 7%), papillary (n = 5, 4%). Two Attention-Based Multiple Instance Learning (ABMIL) classifiers were trained on identical 5-fold cross-validation splits using Virchow2 and concatenated (Virchow2/UNI2/CONCH/GigaPath) embeddings. Prediction uncertainty was quantified via Monte Carlo dropout (T = 30 stochastic forward passes) with entropy of mean probabilities as the uncertainty measure. Multi-model disagreement was computed as a binary indicator when classifier predictions differed. A fixed-weight triage score (0.5× entropy + 0.5×disagreement) ranked cases for referral to pathologist review. Performance was evaluated at 10%, 20%, and 30% referral thresholds on non-referred cases. Results: For five-class LUAD predominant-pattern subtyping, baseline balanced accuracy was 65.2% with macro-F1 of 0.65. Model disagreement strongly predicted classification errors: accuracy was 81.7% when models agreed versus only 44.1% when models disagreed. Uncertainty also tracked errors: entropy was higher in incorrect than correct predictions (0.60 (95% CI 0.51–0.68) vs 0.83 (95% CI 0.71–0.94)). Using the combined uncertainty–disagreement triage score, performance improved on non-referred cases (Table 1). At 30% referral, balanced accuracy increased by 15.6 percentage points. Conclusions: This framework enables risk-stratified AI deployment where the system self-identifies cases requiring expert review. In clinical workflow, the system would prioritize 20–30% of uncertain cases for focused review, while remaining cases would still undergo pathologist sign-out with AI decision support. This addresses a critical barrier to clinical AI adoption: the inability to distinguish reliable from unreliable predictions. The approach transforms AI from an autonomous classifier into an intelligent triage system that appropriately allocates expert pathologist attention to diagnostically challenging cases. Triage performance by referral rat. Referral % Balanced Acc Macro-F1 Cases Retained 0% 65.2% 0.65 143 10% 68.6% 0.66 129 20% 74.4% 0.76 114 30% 80.8% ± 10.1% 0.80 ± 0.09 100
Article Details
Journal Info
Journal of Clinical Oncology
Lippincott Williams & Wilkins
Authors (7)
Meghdad Sabouri Rad
SUNY Upstate Medical University, Syracuse, NY
Mohammad Mehdi Hosseini
SUNY Upstate Medical University, Syracuse, NY
Palak Patel
Polymer Science and Engineering Division, CSIR-National Chemical Laboratory 1 , Pune 411008,
Saverio J. Carello
SUNY Upstate Medical University, Syracuse, NY
Ola El-Zammar
SUNY Upstate Medical University, Syracuse, NY
Michel R. Nasr
3SUNY Upstate University, Department of Pathology, Syracuse, United States
Bardia Rodd
SUNY Upstate Medical University, Syracuse, NY