Beyond convolution: Vision transformer–based modeling for histopathological classification of colorectal carcinoma.
Abstract
e15671 Background: Colorectal carcinoma (CRC) is a major contributor to global cancer morbidity and mortality, with histopathological evaluation remaining the diagnostic gold standard. However, CRC tissue exhibits marked architectural heterogeneity, multi-scale glandular patterns, and variable stromal composition, contributing to diagnostic variability and increased pathologist workload. Convolutional neural networks (CNNs) have demonstrated strong performance in CRC histology classification but are inherently limited by localized receptive fields, which may restrict modeling of long-range spatial relationships critical for complex tissue architectures. Vision Transformers (ViTs) offer an alternative paradigm based on self-attention, enabling global contextual reasoning across entire histopathological images. We evaluated a ViT-B/16 architecture for automated CRC classification using publicly available colorectal histopathology datasets. Methods: Digitized histopathological images from curated colorectal cancer datasets encompassing benign and malignant tissue classes were analyzed. Images underwent preprocessing including patch extraction, stain normalization, and data augmentation. Inputs were resized to 224×224 pixels and partitioned into non-overlapping 16×16 patches. A Vision Transformer B/16 architecture pretrained on ImageNet was fine-tuned for CRC classification. The model incorporated linear patch embeddings, positional encodings, a learnable class token, and 12 transformer encoder layers with multi-head self-attention. Data were split into training and validation cohorts with class balancing applied during training. Model performance was evaluated using accuracy, sensitivity, specificity, precision, recall, F1 score, and confusion matrix analysis. Results: The Vision Transformer achieved high classification performance, with sensitivity of 98% and specificity of 97% across colorectal tissue classes. Global attention-based feature modeling improved discrimination between morphologically similar regions and reduced misclassification in areas with complex glandular architecture or heterogeneous tumor–stroma interfaces. Overall performance was comparable to state-of-the-art CNN-based approaches, with the added benefit of improved global context modeling and enhanced interpretability through attention visualization. Conclusions: Vision Transformer–based modeling provides a robust and complementary approach for colorectal carcinoma histopathological classification by capturing long-range spatial dependencies beyond the capabilities of conventional CNNs. While computationally more intensive, ViT architectures offer unique advantages in modeling complex tissue organization and may serve as valuable components in next-generation AI-assisted diagnostic workflows for colorectal cancer.
Article Details
Journal Info
Journal of Clinical Oncology
Lippincott Williams & Wilkins
Authors (10)
Tanzeela Mariam Shuja
Northwestern McHenry Hospital, Mchenry, IL
Elangovan Krishnan
AIM DOCTOR, Thiruvallur, India, India
Shankar Biswas
Jansi Rani Sethuraj
AIM DOCTOR, Thiruverkadu, India
Kavin Elangovan
AIM DOCTOR, Houston, Texas, United States
Ramya Elangovan
AIM DOCTOR, Houston, Texas, United States
Ekow Pinkrah
1Northwestern McHenry Hospital / Rosalind Franklin University, McHenry, United States
Mevlut Ozmen
1Northwestern McHenry Hospital / Rosalind Franklin University, McHenry, United States
Lavanya Nagarajan
1Northwestern McHenry Hospital / Rosalind Franklin University, McHenry, United States
Aditya Suresh
1Northwestern McHenry Hospital / Rosalind Franklin University, McHenry, United States