Beyond local mucosal patterns: Vision transformer–based modeling for diagnosis of gastric tumors on endoscopic imaging.

K Kavin Elangovan (AIM DOCTOR, Houston, Texas, United States) G Gowrishankar Palaniswamy (8Medical University of South Carolina, Lancaster, United States) E Elangovan Krishnan (AIM DOCTOR, Thiruvallur, India, India) R Ramya Elangovan (AIM DOCTOR, Houston, Texas, United States) M Mustafa Abrar Zaman (AIM DOCTOR, Dhaka, Bangladesh) S Sophia Ahmed J Jansi Rani Sethuraj (AIM DOCTOR, Thiruverkadu, India)

Abstract

e15515 Background: Gastric cancer remains a leading cause of cancer-related mortality worldwide, largely due to delayed diagnosis and subtle early mucosal changes that challenge visual detection. Upper gastrointestinal endoscopy is the primary diagnostic modality for gastric tumors; however, interpretation is highly operator dependent, with significant interobserver variability, particularly for early-stage lesions and flat or diffuse tumors. Convolutional neural networks (CNNs) have demonstrated promise in automated endoscopic image analysis but rely on localized receptive fields that may inadequately capture global mucosal architecture. Vision Transformers (ViTs) introduce a fundamentally different paradigm by leveraging self-attention to model long-range spatial dependencies across entire images. We evaluated a ViT-B/16 model for gastric tumor classification using curated endoscopic imaging datasets. Methods: We analyzed publicly available gastric endoscopy datasets derived from multiple institutions, including malignant gastric tumors, benign gastric lesions, and non-neoplastic mucosa annotated by expert endoscopists with histopathologic confirmation. Images were standardized, augmented, and split into training and validation cohorts using stratified sampling. A Vision Transformer B/16 model pretrained on ImageNet was fine-tuned for multi-class gastric lesion classification. Input images (224×224) were divided into non-overlapping 16×16 patches and embedded into a token sequence augmented with positional encodings and a learnable class token. The architecture employed 12 transformer encoder blocks with multi-head self-attention and feed-forward layers, enabling global contextual reasoning across mucosal patterns. Performance was assessed using accuracy, sensitivity, specificity, F1 score, and area under the receiver operating characteristic curve (AUROC). Results: The Vision Transformer achieved strong diagnostic performance, with overall accuracy exceeding 96% and AUROC greater than 0.96 across gastric lesion categories. Attention-based global modeling improved discrimination of early gastric cancers and lesions with diffuse or irregular mucosal patterns, reducing misclassification commonly observed with convolutional approaches. Performance remained stable across lesion morphology and imaging conditions, supporting generalizability. Conclusions: Vision Transformer–based modeling enables accurate and interpretable diagnosis of gastric tumors by capturing global mucosal and structural context beyond localized feature extraction. While computationally more intensive than CNNs, ViT architectures offer complementary advantages for complex endoscopic imaging tasks and warrant further prospective evaluation to enhance early gastric cancer detection and diagnostic consistency.

Article Details

Volume / Issue Vol. 44, Issue 16_suppl
Published June 01, 2026
ISSN 0732-183X
Publisher Lippincott Williams & Wilkins

Journal Info

Journal of Clinical Oncology

Lippincott Williams & Wilkins

ISSN: 0732-183X Health Sciences

Authors (7)

K

Kavin Elangovan

AIM DOCTOR, Houston, Texas, United States

G

Gowrishankar Palaniswamy

8Medical University of South Carolina, Lancaster, United States

E

Elangovan Krishnan

AIM DOCTOR, Thiruvallur, India, India

R

Ramya Elangovan

AIM DOCTOR, Houston, Texas, United States

M

Mustafa Abrar Zaman

AIM DOCTOR, Dhaka, Bangladesh

S

Sophia Ahmed

J

Jansi Rani Sethuraj

AIM DOCTOR, Thiruverkadu, India