Nominal Compound Chain Extraction Enhanced by Chain-of-Thought Information
摘要
In traditional lexical chain extraction tasks, researchers typically focus on identifying simple lexical items based on surface grammatical relations, often overlooking compound words with underlying semantic frameworks. To address this limitation, the task of Nominal Compound Chain Extraction (NCCE) has emerged. This task aims to identify and cluster nominal compounds sharing the same semantic theme, thereby providing richer semantic information and facilitating a deeper understanding of the latent themes within documents. In this study, we fine-tune the large language model Qwen2-0.5b, employ data augmentation techniques, and introduce Chain-of-Thought (CoT) information from large models as an auxiliary aid, significantly enhancing the model’s document comprehension capabilities.