Multi-Teacher Knowledge Distillation (MKD) improves the performance of the student network by utilizing insights from multiple pre-trained teacher networks, each offering unique types of knowledge. The majority of current MKD approaches concentrate on designing various weighting techniques to form a robust ensemble of teacher networks. However, these methods often overlook the fact that students with poor learning abilities may not derive advantages from such integrated knowledge. To tackle this challenge, we introduce a new method which is Adaptive Multi-Teacher Knowledge Distillation with Class Attention Transfer. Our method supervises students by providing them with an appropriate class attention map from a customized ensemble of teachers. Additionally, by utilizing a meta-weight network, we effectively leverage diverse yet compatible teacher class attention knowledge to improve the performance of the student network. Comprehensive experiments performed on various benchmark datasets demonstrate the effectiveness and adaptability of our method. The results demonstrate significant improvements in the student’s learning outcomes, proving that our adaptive method is better suited to cater to students with varying learning abilities. This tailored approach ensures that even students with initially poor learning capacities can benefit from the knowledge distillation process, ultimately leading to a more robust and capable student network.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Adaptive Multi-teacher Knowledge Distillation with Class Attention Transfer

  • Xin Cheng,
  • Jinjia Zhou

摘要

Multi-Teacher Knowledge Distillation (MKD) improves the performance of the student network by utilizing insights from multiple pre-trained teacher networks, each offering unique types of knowledge. The majority of current MKD approaches concentrate on designing various weighting techniques to form a robust ensemble of teacher networks. However, these methods often overlook the fact that students with poor learning abilities may not derive advantages from such integrated knowledge. To tackle this challenge, we introduce a new method which is Adaptive Multi-Teacher Knowledge Distillation with Class Attention Transfer. Our method supervises students by providing them with an appropriate class attention map from a customized ensemble of teachers. Additionally, by utilizing a meta-weight network, we effectively leverage diverse yet compatible teacher class attention knowledge to improve the performance of the student network. Comprehensive experiments performed on various benchmark datasets demonstrate the effectiveness and adaptability of our method. The results demonstrate significant improvements in the student’s learning outcomes, proving that our adaptive method is better suited to cater to students with varying learning abilities. This tailored approach ensures that even students with initially poor learning capacities can benefit from the knowledge distillation process, ultimately leading to a more robust and capable student network.