The scientific community relies heavily on proper citation practices to acknowledge previous research contributions. Recent efforts of relevant citation identification have focused on using classifiers and feature selection techniques to enhance performance. Utilizing insufficient or imbalanced training data by classifiers may fail to accurately predict relevancy for unseen citations. This work creates dataset and proposes relevance prediction approach (MCTR) for citing and cited articles relying on four factors such as theory congruence rates, citation mention rates, metadata-based, and content-based similarities. The weight distribution of for each factor is determined using Multiple Regression model. Scores are calculated using Cosine Similarity, a Bag-of-Words for Content-based, statistical analysis for Metadata-based, Normal distribution for theory congruence rates and citation mention frequency for mention rates. Experiments conducted on two datasets demonstrate that the proposed approach achieves comparable results of 95% precision on dataset of 465 citation pairs and 98% precision on 500 pairs datasets.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

MCTR–Relevant Citation Identification of Research Articles Based on Metadata, Content, Theory Congruence and Mention Rates Using Multiple Regression

  • Kyu Kyu Win,
  • Yi Yi Hlaing

摘要

The scientific community relies heavily on proper citation practices to acknowledge previous research contributions. Recent efforts of relevant citation identification have focused on using classifiers and feature selection techniques to enhance performance. Utilizing insufficient or imbalanced training data by classifiers may fail to accurately predict relevancy for unseen citations. This work creates dataset and proposes relevance prediction approach (MCTR) for citing and cited articles relying on four factors such as theory congruence rates, citation mention rates, metadata-based, and content-based similarities. The weight distribution of for each factor is determined using Multiple Regression model. Scores are calculated using Cosine Similarity, a Bag-of-Words for Content-based, statistical analysis for Metadata-based, Normal distribution for theory congruence rates and citation mention frequency for mention rates. Experiments conducted on two datasets demonstrate that the proposed approach achieves comparable results of 95% precision on dataset of 465 citation pairs and 98% precision on 500 pairs datasets.