Key phrase extraction techniques are extensively used in computer science, especially in the domains of NLP and information retrieval. They can be applied to text summarization, index building, query improvement, and more. For a wide range of applications, including the rapid and efficient evaluation of large volumes of textual content on the internet, keywords, and key phrases are crucial. A set of representative terms in a document that give readers high-level content specifications are called keywords. This paper describes a Context Specific Key phrase Identification Method (CSKIM), which extracts document key phrases by using prior positive samples of domain key phrases to assign weights to the candidate key phrases. The more keywords a candidate key phrase contains and the more significant these keywords are, the more likely this candidate phrase is a key phrase. CSKIM has the following steps: 1. To obtain prior positive inputs, CSKIM first populates its glossary database using manually identified key phrases and keywords. 2. It then checks the composition of all noun phrases of a document looks up the database and calculates scores for all these noun phrases. The ones having higher scores will be extracted as key phrases 3. Create vector of key phrases.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Enhancing Key Phrase Extraction from Suicidal Data Using CSKIM

  • K. Soumya,
  • Vijay Kumar Garg

摘要

Key phrase extraction techniques are extensively used in computer science, especially in the domains of NLP and information retrieval. They can be applied to text summarization, index building, query improvement, and more. For a wide range of applications, including the rapid and efficient evaluation of large volumes of textual content on the internet, keywords, and key phrases are crucial. A set of representative terms in a document that give readers high-level content specifications are called keywords. This paper describes a Context Specific Key phrase Identification Method (CSKIM), which extracts document key phrases by using prior positive samples of domain key phrases to assign weights to the candidate key phrases. The more keywords a candidate key phrase contains and the more significant these keywords are, the more likely this candidate phrase is a key phrase. CSKIM has the following steps: 1. To obtain prior positive inputs, CSKIM first populates its glossary database using manually identified key phrases and keywords. 2. It then checks the composition of all noun phrases of a document looks up the database and calculates scores for all these noun phrases. The ones having higher scores will be extracted as key phrases 3. Create vector of key phrases.