<p>Existing knowledge distillation methods typically extract rich feature information from data to facilitate knowledge transfer from a large teacher model to a lightweight student network. However, such complex feature representations become highly entangled due to the presence of shared semantic information and overlapping relationships among instance samples. This increases the difficulty of mapping features to labels, ultimately leading to suboptimal performance of the distilled student model. To address this issue, we propose a novel Label-enhanced Contrastive Knowledge Distillation (LCKD) method, which restructures the coupling between features and labels while deeply exploring a complete and reliable knowledge transfer process. Specifically, we design an adaptive label interpolation strategy that dynamically assigns unique virtual labels to data in each training batch. Within our constructed non-overlapping label space, we improve the confidence of pseudo-labels by combining global and local thresholds, creating a smoother mapping landscape. Furthermore, pseudo-labels provide logit-level supervision, while contrastive alignment enforces feature-level consistency. This multi-level process allows the student model to mitigate knowledge loss and ambiguity arising from the teacher’s subjectivity, thereby reducing bias throughout the distillation process. Extensive experiments on image classification tasks validate the effectiveness of the proposed LCKD. Our test results on the CIFAR-100, CIFAR-10, STL-10, Stanford Dogs, Tiny ImageNet and ImageNet datasets surpass those of state-of-the-art (SOTA) knowledge distillation methods. In addition, comprehensive ablation studies also confirm the robustness of our approach.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Label-enhanced contrastive knowledge distillation

  • Zike Qiao,
  • Ze Tao,
  • Lingfeng He,
  • Jian Zhang,
  • Zailiang Chen,
  • Hui Sun

摘要

Existing knowledge distillation methods typically extract rich feature information from data to facilitate knowledge transfer from a large teacher model to a lightweight student network. However, such complex feature representations become highly entangled due to the presence of shared semantic information and overlapping relationships among instance samples. This increases the difficulty of mapping features to labels, ultimately leading to suboptimal performance of the distilled student model. To address this issue, we propose a novel Label-enhanced Contrastive Knowledge Distillation (LCKD) method, which restructures the coupling between features and labels while deeply exploring a complete and reliable knowledge transfer process. Specifically, we design an adaptive label interpolation strategy that dynamically assigns unique virtual labels to data in each training batch. Within our constructed non-overlapping label space, we improve the confidence of pseudo-labels by combining global and local thresholds, creating a smoother mapping landscape. Furthermore, pseudo-labels provide logit-level supervision, while contrastive alignment enforces feature-level consistency. This multi-level process allows the student model to mitigate knowledge loss and ambiguity arising from the teacher’s subjectivity, thereby reducing bias throughout the distillation process. Extensive experiments on image classification tasks validate the effectiveness of the proposed LCKD. Our test results on the CIFAR-100, CIFAR-10, STL-10, Stanford Dogs, Tiny ImageNet and ImageNet datasets surpass those of state-of-the-art (SOTA) knowledge distillation methods. In addition, comprehensive ablation studies also confirm the robustness of our approach.