Distillation of ensemble adaptive teacher assistant knowledge
摘要
Knowledge distillation is a proven compression technique for training student models with the guidance of pre-trained teacher models, highlighting the importance of acquiring various facets of information from these teachers. An alternative method for diversifying knowledge involves incorporating perspectives and characteristics from multiple teachers, which allows students to develop a more nuanced understanding. However, existing multi-teacher distillation methods primarily focus on ensemble strategy design and often overlook the gap between teacher and student models. With this in mind, we propose a distillation methodology named Ensemble of Adaptive Teacher Assistant Knowledge (EATAD), which employs teaching assistants to bridge the gap between teachers and students while offering diverse insights. This method not only enhances knowledge diversity but also facilitates a deeper understanding of teacher behavior among students. Specifically, an adapter module is utilized to fine-tune pre-trained teachers, yielding novel networks termed adaptive teacher assistants (ATA) that help reduce the gap between teachers and students. The diversity of distilled knowledge is improved by simultaneously integrating the knowledge from both teachers and assistants. Experimental results show that the proposed EATAD achieves state-of-the-art performance on standard benchmarks, i.e., CIFAR-100 and ImageNet, under both similar-architecture and cross-architecture settings.