Decoupled Classifier Knowledge Distillation

H Hairui Wang M Mengjie Dong G Guifu Zhu Y Ya Li

Abstract

Mainstream knowledge distillation methods primarily include self-distillation, offline distillation, online distillation, output-based distillation, and feature-based distillation. While each approach has its respective advantages, they are typically employed independently. Simply combining two distillation methods often leads to redundant information. If the information conveyed by both methods is highly similar, this can result in wasted computational resources and increased complexity. To provide a new perspective on distillation research, we aim to explore a compromise solution that aligns complex features without conflicting with output alignment. In this work, we propose to decouple the classifier’s output into two components: non-target classes learned by the student, and target classes obtained by both the teacher and the student. Finally, we introduce Decoupled Classifier Knowledge Distillation (DCKD), where on one hand, we fix the correct knowledge that the student has already acquired, which is crucial for merging the two methods; on the other hand, we encourage the student to further align its output with that of the teacher. Compared to using a single method, DCKD achieves superior results on both the CIFAR-100 and ImageNet datasets for image classification and object detection tasks, without reducing training efficiency. Moreover, it allows relational-based and feature-based distillation to operate more efficiently and flexibly. This work demonstrates the great potential of integrating distillation methods, and we hope it will inspire future research.

Article Details

Journal PLoS ONE
Volume / Issue Vol. 20, Issue 2
Published February 21, 2025
Pages e0314267
ISSN 1932-6203
Publisher Public Library of Science

Journal Info

PLoS ONE

Public Library of Science

ISSN: 1932-6203 Open Access Health Sciences

Authors (4)

H

Hairui Wang

M

Mengjie Dong

G

Guifu Zhu

Y

Ya Li