Scientific discovery entails a detailed understanding and structuring of existing hypotheses—a challenging task due to the variety and complexity of the scientific texts. Despite efforts in domains like bio-medicine and invasion biology, there does not seem to be a curated fine-grained hypothesis dataset derived from scientific articles. This paper presents SciHyp, a novel dataset containing the RDF description of 689 unique hypothesis sentences from 479 scientific articles in the computer science domain. The dataset describes hypotheses of two types: relation-finding hypotheses (526), which indicate a relationship between variables, and comparative (274) hypotheses, which specify comparisons between samples based on a variable and an operator. We created the dataset using a novel and multi-step annotation pipeline incorporating expert annotation, Large Language Models (LLMs) including BERT, Sci-BERT, and crowd-based refinement. Our pipeline effectively identified non-hypothesis sentences with a 96.1% consensus rate between the LLMs and crowd annotations, demonstrating its effectiveness in identifying relevant sentences that contain hypotheses. Furthermore, we extracted the individual components of hypotheses (i.e., their variables and the relation between them) using an in-context learning approach based on GPT-4. We believe the SciHyp dataset will benefit the scientific community by offering a structured dataset for model training and evaluation, and adapting the procedure to curate and analyse large-scale hypothesis datasets.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

SciHyp: A Fine-Grained Dataset Describing Hypotheses and Their Components from Scientific Articles

  • Rosni Vasu,
  • Cristina Sarasua,
  • Abraham Bernstein

摘要

Scientific discovery entails a detailed understanding and structuring of existing hypotheses—a challenging task due to the variety and complexity of the scientific texts. Despite efforts in domains like bio-medicine and invasion biology, there does not seem to be a curated fine-grained hypothesis dataset derived from scientific articles. This paper presents SciHyp, a novel dataset containing the RDF description of 689 unique hypothesis sentences from 479 scientific articles in the computer science domain. The dataset describes hypotheses of two types: relation-finding hypotheses (526), which indicate a relationship between variables, and comparative (274) hypotheses, which specify comparisons between samples based on a variable and an operator. We created the dataset using a novel and multi-step annotation pipeline incorporating expert annotation, Large Language Models (LLMs) including BERT, Sci-BERT, and crowd-based refinement. Our pipeline effectively identified non-hypothesis sentences with a 96.1% consensus rate between the LLMs and crowd annotations, demonstrating its effectiveness in identifying relevant sentences that contain hypotheses. Furthermore, we extracted the individual components of hypotheses (i.e., their variables and the relation between them) using an in-context learning approach based on GPT-4. We believe the SciHyp dataset will benefit the scientific community by offering a structured dataset for model training and evaluation, and adapting the procedure to curate and analyse large-scale hypothesis datasets.