Beyond local mucosal patterns: Vision transformer–based modeling for diagnosis of gastric tumors on endoscopic imaging.
Abstract
e15515 Background: Gastric cancer remains a leading cause of cancer-related mortality worldwide, largely due to delayed diagnosis and subtle early mucosal changes that challenge visual detection. Upper gastrointestinal endoscopy is the primary diagnostic modality for gastric tumors; however, interpretation is highly operator dependent, with significant interobserver variability, particularly for early-stage lesions and flat or diffuse tumors. Convolutional neural networks (CNNs) have demonstrated promise in automated endoscopic image analysis but rely on localized receptive fields that may inadequately capture global mucosal architecture. Vision Transformers (ViTs) introduce a fundamentally different paradigm by leveraging self-attention to model long-range spatial dependencies across entire images. We evaluated a ViT-B/16 model for gastric tumor classification using curated endoscopic imaging datasets. Methods: We analyzed publicly available gastric endoscopy datasets derived from multiple institutions, including malignant gastric tumors, benign gastric lesions, and non-neoplastic mucosa annotated by expert endoscopists with histopathologic confirmation. Images were standardized, augmented, and split into training and validation cohorts using stratified sampling. A Vision Transformer B/16 model pretrained on ImageNet was fine-tuned for multi-class gastric lesion classification. Input images (224×224) were divided into non-overlapping 16×16 patches and embedded into a token sequence augmented with positional encodings and a learnable class token. The architecture employed 12 transformer encoder blocks with multi-head self-attention and feed-forward layers, enabling global contextual reasoning across mucosal patterns. Performance was assessed using accuracy, sensitivity, specificity, F1 score, and area under the receiver operating characteristic curve (AUROC). Results: The Vision Transformer achieved strong diagnostic performance, with overall accuracy exceeding 96% and AUROC greater than 0.96 across gastric lesion categories. Attention-based global modeling improved discrimination of early gastric cancers and lesions with diffuse or irregular mucosal patterns, reducing misclassification commonly observed with convolutional approaches. Performance remained stable across lesion morphology and imaging conditions, supporting generalizability. Conclusions: Vision Transformer–based modeling enables accurate and interpretable diagnosis of gastric tumors by capturing global mucosal and structural context beyond localized feature extraction. While computationally more intensive than CNNs, ViT architectures offer complementary advantages for complex endoscopic imaging tasks and warrant further prospective evaluation to enhance early gastric cancer detection and diagnostic consistency.
Article Details
Journal Info
Journal of Clinical Oncology
Lippincott Williams & Wilkins
Authors (7)
Kavin Elangovan
AIM DOCTOR, Houston, Texas, United States
Gowrishankar Palaniswamy
8Medical University of South Carolina, Lancaster, United States
Elangovan Krishnan
AIM DOCTOR, Thiruvallur, India, India
Ramya Elangovan
AIM DOCTOR, Houston, Texas, United States
Mustafa Abrar Zaman
AIM DOCTOR, Dhaka, Bangladesh
Sophia Ahmed
Jansi Rani Sethuraj
AIM DOCTOR, Thiruverkadu, India