Open Information Extraction (OIE) is an important subtask of information extraction that freely extracts triples with subject-relation-object structures from text. Sequence labeling methods often struggle to produce outputs beyond the original sequence, whereas sequence generation methods are more effective in handling this task. Sequence generation methods typically model OIE as a text-to-structure generation task. However, common sequence generation methods exhibit deficiencies in semantic understanding, which can lead to incomplete information in single extraction and error information in multiple extractions. Therefore, we propose a two-stage semantic enhancement method for open information extraction, which enhances semantic understanding during the training stage and introduces semantic constraints during the inference stage. Specifically, in the training stage, we incorporate three auxiliary tasks to improve the model's semantic understanding; in the inference stage, we introduce a multi-granularity iterative reranker to impose semantic constraints to capture results that are more semantically aligned. Experimental results demonstrate that our approach achieves comparable performance on the CaRB dataset and leads the field among generation-based methods.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Two-Stage Semantic Enhancement for Open Information Extraction

  • Yi Sheng,
  • Huizhe Su,
  • Xiangfeng Luo,
  • Xinzhi Wang,
  • Shaorong Xie

摘要

Open Information Extraction (OIE) is an important subtask of information extraction that freely extracts triples with subject-relation-object structures from text. Sequence labeling methods often struggle to produce outputs beyond the original sequence, whereas sequence generation methods are more effective in handling this task. Sequence generation methods typically model OIE as a text-to-structure generation task. However, common sequence generation methods exhibit deficiencies in semantic understanding, which can lead to incomplete information in single extraction and error information in multiple extractions. Therefore, we propose a two-stage semantic enhancement method for open information extraction, which enhances semantic understanding during the training stage and introduces semantic constraints during the inference stage. Specifically, in the training stage, we incorporate three auxiliary tasks to improve the model's semantic understanding; in the inference stage, we introduce a multi-granularity iterative reranker to impose semantic constraints to capture results that are more semantically aligned. Experimental results demonstrate that our approach achieves comparable performance on the CaRB dataset and leads the field among generation-based methods.