This chapter addresses the foundational challenge of constructing student-teacher architectures in KD. Traditional KD frameworks rely on pre-trained teacher models to guide smaller student networks, but face two critical limitations: (1) the two-stage training process incurs excessive computational overhead, particularly prohibitive for modern large-scale models, and (2) performance instability arises from suboptimal teacher selection, where high-accuracy teachers may paradoxically degrade student performance. To solve these issues, we propose self-distillation as an innovative paradigm that eliminates external teacher dependency. This approach introduces shallow classifiers at intermediate layers to form a multi-exit architecture, enabling intra-model knowledge transfer through hierarchical layer-wise supervision. These findings position self-distillation as a unified solution addressing computational efficiency, deployment flexibility, and performance stability in KD frameworks.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Student and Teacher Models in KD

  • Linfeng Zhang

摘要

This chapter addresses the foundational challenge of constructing student-teacher architectures in KD. Traditional KD frameworks rely on pre-trained teacher models to guide smaller student networks, but face two critical limitations: (1) the two-stage training process incurs excessive computational overhead, particularly prohibitive for modern large-scale models, and (2) performance instability arises from suboptimal teacher selection, where high-accuracy teachers may paradoxically degrade student performance. To solve these issues, we propose self-distillation as an innovative paradigm that eliminates external teacher dependency. This approach introduces shallow classifiers at intermediate layers to form a multi-exit architecture, enabling intra-model knowledge transfer through hierarchical layer-wise supervision. These findings position self-distillation as a unified solution addressing computational efficiency, deployment flexibility, and performance stability in KD frameworks.