Multi-label Short Text Classification (MSTC) focuses on assigning multiple relevant labels to each short text. However, it is confronted with several challenges such as data sparsity, label scarcity, and class imbalance. Motivated by this, in this paper, we introduce a novel MSTC method called Knowledge-based Long-tail Data Augmentation and Pseudo-label optimization (KLDAP). Specifically, to fix the data sparsity, KLDAP firstly enriches the feature representation of short texts by integrating external knowledge from concept knowledge graphs and entity recognition techniques. Secondly, it incorporates an optimized semi-supervised learning framework to enhance the utilization of unlabeled data with the help of external knowledge and generated pseudo labels. Thirdly, to further address the class imbalance, we introduce an innovative data augmentation strategy that combines a Variational Autoencoder (VAE) with head label feature transfer and contrastive learning, optimizing the representation of tail labels. Finally, extensive experiments conducted on four benchmark datasets under varying labeling ratios demonstrate that KLDAP significantly outperforms state-of-the-art methods, effectively tackling the challenges in MSTC.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

External Knowledge-Enhanced Semi-Supervised Multi-Label Short Text Classification

  • Zhi-Jie Wang,
  • Yirui Li,
  • Shuhui Cao,
  • Peipei Li

摘要

Multi-label Short Text Classification (MSTC) focuses on assigning multiple relevant labels to each short text. However, it is confronted with several challenges such as data sparsity, label scarcity, and class imbalance. Motivated by this, in this paper, we introduce a novel MSTC method called Knowledge-based Long-tail Data Augmentation and Pseudo-label optimization (KLDAP). Specifically, to fix the data sparsity, KLDAP firstly enriches the feature representation of short texts by integrating external knowledge from concept knowledge graphs and entity recognition techniques. Secondly, it incorporates an optimized semi-supervised learning framework to enhance the utilization of unlabeled data with the help of external knowledge and generated pseudo labels. Thirdly, to further address the class imbalance, we introduce an innovative data augmentation strategy that combines a Variational Autoencoder (VAE) with head label feature transfer and contrastive learning, optimizing the representation of tail labels. Finally, extensive experiments conducted on four benchmark datasets under varying labeling ratios demonstrate that KLDAP significantly outperforms state-of-the-art methods, effectively tackling the challenges in MSTC.