Study on Extracting Keywords that Reveal the Value of Research Data Through Comparisons Between Academic and Data Papers
摘要
The objective of this study is to identify keywords that are considered to be metadata and that influence the utilization of research data in repositories. In particular, a comparative analysis was conducted of the linguistic characteristics of a data paper, which differ from those typically found in academic papers. The extraction of keywords that represent these features could become a crucial aspect of metadata, facilitating the use of research data. We prepared a data paper and an academic paper authored by the same researchers, both addressing related topics. Additionally, we collected eight academic papers that cited the data paper. Using natural language processing techniques, we extracted the distinctive terms unique to the data paper. Furthermore, we analyzed the terms, and the terms not present in the titles or abstracts of the data paper may serve as keywords that represent data characteristics, having a significant impact on the citation of data papers.