<p>The accuracy of bibliometric databases in classifying <i>document types</i> (DTs)—such as <i>research articles</i>, <i>conference proceedings</i>, <i>reviews</i>, <i>short notes</i>, <i>letters</i>, <i>book chapters</i>, etc.—is crucial for the academic community, as bibliometric indicators may significantly influence research funding, decision-making, and academic reputation. This study presents a semi-automated methodology to assess the accuracy of DT classification in bibliometric databases, such as Scopus and Web of Science (WoS). The methodology can handle large document volumes and adapt to different DT categories without predefined correspondences. The first phase of the methodology automatically identifies discrepancies in DT classifications between Scopus and WoS, in order to find potentially misclassified documents; the second phase involves manually analyzing these documents to confirm and attribute classification errors. The methodology is applied to a sample of several tens of thousands of papers from the teaching staff of two major universities in Turin (Italy). The results show overall error rates of approximately 2.7% for Scopus and 2.3% for WoS. The paper also analyzes the most common types of errors found in both databases, providing an interpretation of these inaccuracies and some insights for possible improvements in the quality of these databases.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A large-scale semi-automated approach for assessing document-type classification errors in bibliometric databases

  • D. A. Maisano,
  • L. Mastrogiacomo,
  • L. Ferrara,
  • F. Franceschini

摘要

The accuracy of bibliometric databases in classifying document types (DTs)—such as research articles, conference proceedings, reviews, short notes, letters, book chapters, etc.—is crucial for the academic community, as bibliometric indicators may significantly influence research funding, decision-making, and academic reputation. This study presents a semi-automated methodology to assess the accuracy of DT classification in bibliometric databases, such as Scopus and Web of Science (WoS). The methodology can handle large document volumes and adapt to different DT categories without predefined correspondences. The first phase of the methodology automatically identifies discrepancies in DT classifications between Scopus and WoS, in order to find potentially misclassified documents; the second phase involves manually analyzing these documents to confirm and attribute classification errors. The methodology is applied to a sample of several tens of thousands of papers from the teaching staff of two major universities in Turin (Italy). The results show overall error rates of approximately 2.7% for Scopus and 2.3% for WoS. The paper also analyzes the most common types of errors found in both databases, providing an interpretation of these inaccuracies and some insights for possible improvements in the quality of these databases.