Knowledge distillation, a popular model compression method, transfers knowledge from a large teacher model to a smaller student model. Self-distillation takes this a step further by having the model itself act as both teacher and student. However, existing self-distillation methods often focus on individual instance knowledge, such as logits and intermediate features, but overlook the structural information within each category’s representation. To address this gap, we propose Self-Distillation via Intra-Class Compactness (SDICC). Specifically, in SDICC, we use previous epoch models as teachers to guide training in the current epoch, while also emphasizing intra-class compactness as an additional training objective. This facilitates our model’s learning process in bringing intra-class features closer together, thereby promoting more discriminative representations across different categories. Moreover, to better combine both the knowledge from logits and the compactness of features, we adaptively perform self-distillation for progressive knowledge transfer. We extensively evaluate SDICC on popular image classification datasets like CIFAR-100 and Tiny ImageNet. Our results demonstrate that SDICC outperforms recent state-of-the-art self-distillation methods, showcasing its effectiveness in knowledge transfer and model compression.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Self-Distillation via Intra-Class Compactness

  • Jiaye Lin,
  • Lin Li,
  • Baosheng Yu,
  • Weihua Ou,
  • Jianping Gou

摘要

Knowledge distillation, a popular model compression method, transfers knowledge from a large teacher model to a smaller student model. Self-distillation takes this a step further by having the model itself act as both teacher and student. However, existing self-distillation methods often focus on individual instance knowledge, such as logits and intermediate features, but overlook the structural information within each category’s representation. To address this gap, we propose Self-Distillation via Intra-Class Compactness (SDICC). Specifically, in SDICC, we use previous epoch models as teachers to guide training in the current epoch, while also emphasizing intra-class compactness as an additional training objective. This facilitates our model’s learning process in bringing intra-class features closer together, thereby promoting more discriminative representations across different categories. Moreover, to better combine both the knowledge from logits and the compactness of features, we adaptively perform self-distillation for progressive knowledge transfer. We extensively evaluate SDICC on popular image classification datasets like CIFAR-100 and Tiny ImageNet. Our results demonstrate that SDICC outperforms recent state-of-the-art self-distillation methods, showcasing its effectiveness in knowledge transfer and model compression.