Keyphrases represent the core content of a document, facilitating efficient information processing in knowledge discovery systems. Traditional supervised keyphrase prediction methods not only rely on labeled data but also lack robustness across diverse domains. Recently, with the advancements in pretrained language models, unsupervised methods have gained increasing attention. This survey provides a comprehensive review of the entire process of unsupervised keyphrase prediction. We begin with an analysis of the linguistic properties of keyphrases, aiming to support the design and evaluation of keyphrase prediction models. Then, we categorize and discuss the details of existing unsupervised methods for both keyphrase extraction and generation, emphasizing cutting-edge techniques such as attention mechanisms and prompt learning. Additionally, we examine evaluation metrics, introduce a novel reference-free metric, and provide a list of open-source datasets. Finally, we explore promising future directions and conclude the survey.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Unsupervised Keyphrase Prediction: Methods and Evaluation

  • Huiqian Wu,
  • Yuqing Sun

摘要

Keyphrases represent the core content of a document, facilitating efficient information processing in knowledge discovery systems. Traditional supervised keyphrase prediction methods not only rely on labeled data but also lack robustness across diverse domains. Recently, with the advancements in pretrained language models, unsupervised methods have gained increasing attention. This survey provides a comprehensive review of the entire process of unsupervised keyphrase prediction. We begin with an analysis of the linguistic properties of keyphrases, aiming to support the design and evaluation of keyphrase prediction models. Then, we categorize and discuss the details of existing unsupervised methods for both keyphrase extraction and generation, emphasizing cutting-edge techniques such as attention mechanisms and prompt learning. Additionally, we examine evaluation metrics, introduce a novel reference-free metric, and provide a list of open-source datasets. Finally, we explore promising future directions and conclude the survey.