In the recent years, Transformers have emerged as some of the most powerful models for sequential data processing, demonstrating exceptional performance across various domains. Their ability to capture long-term dependencies is remarkable and has given profound inference and detection capabilities. However, their computational demands for training and inference pose significant challenges. This study introduces two optimization techniques for transformers: uncertainty sampling-based active learning to enhance robustness and accuracy, and student-teacher knowledge distillation to improve efficiency. Experimental results on the Internet Movie databases dataset indicate a 7 \(\times \) speedup in inference time with an acceptable trade-off in accuracy, showcasing the effectiveness of our approach.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Enhancing Transformer Efficiency Through Active Learning and Knowledge Distillation

  • Dhruv Kulkarni,
  • Yagnik Dhameliya,
  • Samarth Gupta,
  • Chandra Prakash

摘要

In the recent years, Transformers have emerged as some of the most powerful models for sequential data processing, demonstrating exceptional performance across various domains. Their ability to capture long-term dependencies is remarkable and has given profound inference and detection capabilities. However, their computational demands for training and inference pose significant challenges. This study introduces two optimization techniques for transformers: uncertainty sampling-based active learning to enhance robustness and accuracy, and student-teacher knowledge distillation to improve efficiency. Experimental results on the Internet Movie databases dataset indicate a 7 \(\times \) speedup in inference time with an acceptable trade-off in accuracy, showcasing the effectiveness of our approach.