Distilled Knowledge in KD
摘要
This chapter systematically discusses the second fundamental problem in KD—what knowledge to be distilled. While traditional KD focuses on logits (soft predictions), subsequent research reveals richer information in feature maps and relational patterns. However, the optimal knowledge type remains unclear, as neural networks lack explicit metrics to evaluate knowledge utility. This chapter studies the influence from the distilled knowledge in KD, from the perspectives of both task-oriented KD and task-irrelevant KD. Our findings demonstrate that task-oriented KD is able to transfer the most crucial knowledge for the given task, while task-irrelevant KD is more beneficial in all kinds of downstream tasks, such as classification, detection, and segmentation on images and videos.