Research data are cited in scholarly papers, and the construction and use of datasets are mentioned. The descriptions of research data in papers may be used as information for its metadata. In this paper, we focus on large language models (LLM), which have achieved high performance in various natural language processing tasks, and we investigate the ability of LLMs to extract metadata from papers. In the experiment, we analyzed LLMs’ metadata extraction capabilities quantitatively and qualitatively. The results demonstrate that while LLMs can extract metadata from papers extensively, the extraction accuracy is not necessarily high. We confirm that there are challenges in identifying the names of research data and linking information related to the research data.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Capabilities and Challenges of LLMs in Metadata Extraction from Scholarly Papers

  • Yu Watanabe,
  • Koichiro Ito,
  • Shigeki Matsubara

摘要

Research data are cited in scholarly papers, and the construction and use of datasets are mentioned. The descriptions of research data in papers may be used as information for its metadata. In this paper, we focus on large language models (LLM), which have achieved high performance in various natural language processing tasks, and we investigate the ability of LLMs to extract metadata from papers. In the experiment, we analyzed LLMs’ metadata extraction capabilities quantitatively and qualitatively. The results demonstrate that while LLMs can extract metadata from papers extensively, the extraction accuracy is not necessarily high. We confirm that there are challenges in identifying the names of research data and linking information related to the research data.