Deep learning–based classification of benign and malignant breast lesions on ultrasound using knowledge distillation with external validation and global deployment.
Abstract
556 Background: Breast cancer is the most common malignancy in women worldwide. Although breast ultrasonography is widely accessible, diagnostic accuracy remains limited by operator dependence and interobserver variability (κ≈0.6–0.8), contributing to delayed diagnosis and unnecessary biopsies. While deep learning has shown promise in ultrasound-based lesion classification, many models are computationally intensive and lack external validation, limiting clinical scalability. We developed a deployment-ready deep learning framework that balances diagnostic performance with computational efficiency using structured knowledge distillation. Objectives: To develop and externally validate a computationally efficient deep learning model for binary breast ultrasound lesion classification; to evaluate whether a distilled student model preserves clinically meaningful performance from a high-capacity teacher network; and to assess global clinical feasibility through independent radiologist evaluation. Methods: We analyzed 8,116 breast ultrasound images (4,074 benign; 4,042 malignant) independently annotated by two board-certified radiologists with consensus adjudication. A ResNet34 teacher network (21.8M parameters) was trained using ImageNet-pretrained weights. A lightweight ResNet18 student network (11.7M parameters; 46% reduction) was trained using structured knowledge distillation with confidence-aligned soft targets and KL-divergence regularization. External validation was performed on three independent ultrasound datasets not used during training. Performance metrics included accuracy, sensitivity, specificity, precision, F1 score, and AUROC. The model was deployed in a HIPAA-compliant cross-platform application and independently evaluated by physicians across six continents. Results: The student model achieved 91.9% accuracy on held-out test data with balanced performance (sensitivity 90.8%, specificity 93.0%, F1 score 0.918). External validation demonstrated stable generalization (accuracy 89–92%). In global clinical deployment, 94.2% of participating physicians rated the system clinically useful and suitable for integration into routine breast ultrasound workflows. Conclusions: Knowledge distillation enables substantial model compression while preserving diagnostic accuracy for breast lesion classification on ultrasound. This scalable, deployment-ready approach supports standardized interpretation across diverse clinical environments and warrants prospective evaluation for reducing unnecessary biopsies and improving diagnostic efficiency.
Article Details
Journal Info
Journal of Clinical Oncology
Lippincott Williams & Wilkins
Authors (8)
Jansi Rani Sethuraj
AIM DOCTOR, Thiruverkadu, India
Elangovan Krishnan
AIM DOCTOR, Thiruvallur, India, India
Gowrishankar Palaniswamy
8Medical University of South Carolina, Lancaster, United States
Sravani Bhavanam
2Brookdale University Hospital and Medical center, Brooklyn, United States
Sophia Ahmed
Kavin Elangovan
AIM DOCTOR, Houston, Texas, United States
Ramya Elangovan
AIM DOCTOR, Houston, Texas, United States
Hammad Khan
Ayub Medical College, Abbottabad, Pakistan