A compact texture frequency spatial gating network for oral cancer classification from clinical oral images
Abstract
Abstract Image-level classification of oral-cavity photographs remains challenging because oral-cancer appearances are heterogeneous and many public datasets lack patient identifiers, acquisition metadata, biopsy confirmation, and lesion-level annotations. This study proposes OCAT-Net-S, a compact hierarchical network for binary classification of Normal versus Oral Cancer from RGB clinical oral images. The model combines a convolutional stem, MBConv blocks, MBConv-SE blocks, and Texture Frequency Spatial Gating (TFSG) blocks to integrate local texture extraction, channel recalibration, and frequency-guided spatial attention within a 3.55 M-parameter architecture. Internal evaluation used the Atef-Kaggle Oral Cancer Images for Classification dataset, containing 1,238 images: 553 Normal and 685 Oral Cancer. Because patient identifiers were unavailable, five-fold stratified group cross-validation with filename-sequential grouping was used as a leakage-reduction proxy and repeated across five seeds. OCAT-Net-S achieved an internal AUC-ROC of 0.956, oral-cancer F1-score of 0.942, and sensitivity of 0.942, compared with AUC-ROC 0.941 for the strongest evaluated baseline, TinyViT-5 M. Validation-only temperature scaling produced a mean Brier score of 0.0419, ECE-15 of 0.0174, and NLL of 0.1798. To reduce overinterpretation from repeated seed runs, fold-level inference was reported, giving an internal AUC-ROC estimate of 0.956 with an approximate 95% CI of 0.940–0.972. External transfer testing was performed without external threshold tuning on two independent settings: SMART-OM Normal-versus-OSCC and the Oral Images Dataset benign-versus-malignant oral-lesion task. AUC-ROC decreased to 0.830 and 0.840, respectively, indicating measurable cross-dataset shift despite partial transferability. Qualitative Grad-CAM visualizations suggested lesion-focused activation patterns, but lesion-level masks were unavailable for quantitative localization validation. Overall, OCAT-Net-S demonstrates compact internal image-level discrimination and limited external transferability; however, the findings do not establish patient-level generalization, prospective clinical validity, clinician-equivalent performance, or deployment readiness.
Article Details
Authors (8)
Mahmudul Haque Rijvi
Preety Arifa Momotaj
Syed Mohammed Muhive Uddin
Udoy Sankar Saha
Sazzat Hossain
Md Asikur Rahman Chy
Mia Md Tofayel Gonee Manik
Abdul Basit