Knowledge concept tagging aims to tag exercises with specific knowledge concepts. Traditional manual annotation methods are increasingly unable to meet the demands of annotating large-scale, high-quality data. While automated methods have been developed to streamline this process, they still struggle with fine-grained knowledge concept labeling. This limitation arises from the inherent difficulty of accurately annotating text content with specific knowledge concepts selected from a vast pool of candidate concepts, and the similarity of fine-grained knowledge concepts further exacerbates this challenge. In this paper, we propose a Local-Global Cascaded Ensemble Learning (LGCEL) method based on hybrid experts for knowledge concept tagging to address this issue. Specifically, LGCEL first conducts unsupervised domain continual pre-training, e.g., masked language modeling and causal language modeling, to obtain in-domain models, then performs supervised fine-tuning to achieve diverse tagging experts, and finally collaborates hybrid fine-trained models including lightweight large language models (LLMs) for voting ensembling. Experimental results demonstrate the effectiveness of LGCEL in annotating fine-grained knowledge concepts and its superiority in integrating mixed tagging experts to improve the annotation accuracy of hard-to-distinguish knowledge concepts.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Local-Global Cascaded Ensemble Learning on Hybrid Experts for Knowledge Concept Tagging

  • Zhiwei Yang,
  • Jiahua Yang,
  • Longtao Wang,
  • Rongxin Huo,
  • Huiru Lin,
  • Yuxuan Zhou,
  • Weiqi Luo

摘要

Knowledge concept tagging aims to tag exercises with specific knowledge concepts. Traditional manual annotation methods are increasingly unable to meet the demands of annotating large-scale, high-quality data. While automated methods have been developed to streamline this process, they still struggle with fine-grained knowledge concept labeling. This limitation arises from the inherent difficulty of accurately annotating text content with specific knowledge concepts selected from a vast pool of candidate concepts, and the similarity of fine-grained knowledge concepts further exacerbates this challenge. In this paper, we propose a Local-Global Cascaded Ensemble Learning (LGCEL) method based on hybrid experts for knowledge concept tagging to address this issue. Specifically, LGCEL first conducts unsupervised domain continual pre-training, e.g., masked language modeling and causal language modeling, to obtain in-domain models, then performs supervised fine-tuning to achieve diverse tagging experts, and finally collaborates hybrid fine-trained models including lightweight large language models (LLMs) for voting ensembling. Experimental results demonstrate the effectiveness of LGCEL in annotating fine-grained knowledge concepts and its superiority in integrating mixed tagging experts to improve the annotation accuracy of hard-to-distinguish knowledge concepts.