Abstract <p>The metadata of scientific publications are used to build catalogs, determine their citation, and perform other tasks. The automation of metadata extraction from PDF files provides a means of speeding up the execution of the designated tasks, while the possibility of further use of the obtained data depends on the quality of the extraction. Existing software solutions were analyzed, and three were selected: GROBID, CERMINE, and ScientificPdfParser. A procedure for comparing software solutions to recognize texts of scientific publications according to the quality of the metadata extraction is proposed. Based on this procedure, an experiment was conducted to extract four types of metadata (title, abstract, publication date, and author names). To compare software solutions, a dataset of 112 457 publications divided into 23 subject areas formed on the basis of Semantic Scholar data was used. An example of choosing an effective software solution for metadata extraction under the conditions of specified priorities for subject areas and types of metadata using a weighted sum is given. It was determined that for the given example, CERMINE shows efficiency 10.5% higher than GROBID and 9.6% higher than ScientificPdfParser.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Procedure for Comparing Text Recognition Software Solutions for Scientific Publications by the Quality of Metadata Extraction

  • I. I. Kuznetsov,
  • O. P. Novikov,
  • D. Y. Ilin

摘要

Abstract

The metadata of scientific publications are used to build catalogs, determine their citation, and perform other tasks. The automation of metadata extraction from PDF files provides a means of speeding up the execution of the designated tasks, while the possibility of further use of the obtained data depends on the quality of the extraction. Existing software solutions were analyzed, and three were selected: GROBID, CERMINE, and ScientificPdfParser. A procedure for comparing software solutions to recognize texts of scientific publications according to the quality of the metadata extraction is proposed. Based on this procedure, an experiment was conducted to extract four types of metadata (title, abstract, publication date, and author names). To compare software solutions, a dataset of 112 457 publications divided into 23 subject areas formed on the basis of Semantic Scholar data was used. An example of choosing an effective software solution for metadata extraction under the conditions of specified priorities for subject areas and types of metadata using a weighted sum is given. It was determined that for the given example, CERMINE shows efficiency 10.5% higher than GROBID and 9.6% higher than ScientificPdfParser.