Developing a FAIR data policy for chemical risk assessment and exploring future topics and stakeholder interests with large language models and data sciences
摘要
The EU Partnership for the Assessment of Risk from Chemicals (PARC) develops novel methods for human health and environmental risk assessment (RA) for regulatory application. Large amounts of environmental monitoring, in vitro, in silico and other types of data are generated. To facilitate data reuse and information exchange within and between research and regulatory stakeholders, it is important to document and manage data. In 2016, the FAIR (Findable, Accessible, Interoperable and Reusable) data principles were introduced to foster reuse of data. To adopt FAIR within PARC, a PARC FAIR data policy (PFDP) was developed.
ResultsTo determine PFDP topics, expert panel discussions were held, followed by review of 37 stakeholder documents issued by authorities in RA and environmental, health, (open) data sciences. This resulted in a taxonomy of 57 topics (“narrower term(s)” (NT(s), together with 15 “broader terms”), forming the basis of the PFDP. These included, aside FAIR, aspects specifically relevant for RA, such as GDPR, sensitive data, transparency and data valuation. Large Language Models (LLMs) were used to contrive novel terms for PFDP revisions. Seventy-eight novel terms were discovered by LLMs, compared to 57 NTs already identified by experts. Based on a frequency analysis within stakeholder documents, of all 135 terms (57 + 78) -further categorized under eight ChatGPT-derived “FAIR categories”- and subsequent cluster analyses (PCA, K-means), differences between stakeholders were identified. Organizations concerned with environmental and human chemical RA (e.g. European Chemicals Agency, European Food Safety Authority) could be distinguished from those focused on technical FAIR data issues and those involved in publicly funded research and open science.
ConclusionsTo foster environmental and human health research data for research purposes and adoption of research data in regulatory applications, a taxonomy of terms and FAIR data policy was developed. A frequency analysis of the taxonomy (augmented with LLM-derived) terms in stakeholder documents revealed differences between stakeholders, which may help identify trends e.g. the adoption of FAIR principles within and across sciences and regulatory domains for (NG) RA. The PFDP is designed to support data driven RA improvement within PARC but can equally contribute to FAIR data in other environmental RA research and regulatory communities.