In national statistical offices and central banks, research data centers (RDCs) are tasked with providing access to granular administrative data and have proliferated in recent years. RDCs face challenges in tracing usage of their data in research papers: Tracing usage in papers to date relies on human readers and therefore remains time-consuming and error-prone. To address this, we explore the potential of using large language models (LLMs), specifically GPT-3.5, to automate the identification of data sources. Based on a comprehensive sample of research papers, we create a human-labeled validation dataset and analyze the accuracy of GPT-3.5 in detecting and summarizing data sources in economics and finance papers. Furthermore, we evaluate the detection and prediction accuracy and address the issue of false answers provided by the model. We find that LLMs can advance the status quo considerably: Our results indicate that LLMs can accurately identify dataset mentions verbatim in up to 62% of cases. When pairing the evaluation of the model with domain knowledge of the sought entities, i.e., allowing for the identification of synonymous dataset names, the model identifies 100% of these entities. Thus, we show that using LLMs to find unstructured dataset mentions can provide a valuable pathway for RDCs and data-providing institutions to gauge the impact of their data provision efforts. Our paper provides a detailed description of the pipeline to implement our solution at data-providing institutions, enabling the understanding of data impact and therefore efficient data provision services.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Building a Retrieval-Augmented Generation Pipeline to Trace Administrative Data Use in Academic Papers

  • Sebastian Seltmann,
  • Emily Kormanyos,
  • Hendrik Christian Doll

摘要

In national statistical offices and central banks, research data centers (RDCs) are tasked with providing access to granular administrative data and have proliferated in recent years. RDCs face challenges in tracing usage of their data in research papers: Tracing usage in papers to date relies on human readers and therefore remains time-consuming and error-prone. To address this, we explore the potential of using large language models (LLMs), specifically GPT-3.5, to automate the identification of data sources. Based on a comprehensive sample of research papers, we create a human-labeled validation dataset and analyze the accuracy of GPT-3.5 in detecting and summarizing data sources in economics and finance papers. Furthermore, we evaluate the detection and prediction accuracy and address the issue of false answers provided by the model. We find that LLMs can advance the status quo considerably: Our results indicate that LLMs can accurately identify dataset mentions verbatim in up to 62% of cases. When pairing the evaluation of the model with domain knowledge of the sought entities, i.e., allowing for the identification of synonymous dataset names, the model identifies 100% of these entities. Thus, we show that using LLMs to find unstructured dataset mentions can provide a valuable pathway for RDCs and data-providing institutions to gauge the impact of their data provision efforts. Our paper provides a detailed description of the pipeline to implement our solution at data-providing institutions, enabling the understanding of data impact and therefore efficient data provision services.