Semantic relations between entities such as places, people, and concepts are a widely used method for representing knowledge. Current approaches to relation extraction typically rely on Large Language Models (LLMs) including BERT. PromptORE (Prompt-based Open Relation Extraction) was developed to enhance relation extraction using LLMs for general-purpose documents. However, it is less effective when applied to historical texts, particularly in languages other than English. In this paper, we present an adaptation of PromptORE for specialized documents, specifically digital transcripts of trials from the Spanish Inquisition. Our approach fine-tunes transformer models with their pre-training objective on the data for inference, a process we call “biasing.” This addresses challenges such as intricate entity arrangements and gender-related issues in Spanish texts. We address these challenges through prompt engineering. Our method is evaluated using encoder-based models and validated through expert assessments. In addition, we use a binary classification to assess performance. Our results demonstrate significant accuracy improvements, with Biased PromptORE models outperforming baseline PromptORE models by up to 50%.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Biased PromptORE: Enhancing Relation Extraction in Gendered Languages and Complex Texts The Case of Spanish Documents from the XVI \(^{{\textbf {th}}}\) Century

  • Michel Boeglin,
  • David Kahn,
  • Héctor López Hidalgo,
  • Josiane Mothe,
  • Diego Ortiz,
  • David Panzoli

摘要

Semantic relations between entities such as places, people, and concepts are a widely used method for representing knowledge. Current approaches to relation extraction typically rely on Large Language Models (LLMs) including BERT. PromptORE (Prompt-based Open Relation Extraction) was developed to enhance relation extraction using LLMs for general-purpose documents. However, it is less effective when applied to historical texts, particularly in languages other than English. In this paper, we present an adaptation of PromptORE for specialized documents, specifically digital transcripts of trials from the Spanish Inquisition. Our approach fine-tunes transformer models with their pre-training objective on the data for inference, a process we call “biasing.” This addresses challenges such as intricate entity arrangements and gender-related issues in Spanish texts. We address these challenges through prompt engineering. Our method is evaluated using encoder-based models and validated through expert assessments. In addition, we use a binary classification to assess performance. Our results demonstrate significant accuracy improvements, with Biased PromptORE models outperforming baseline PromptORE models by up to 50%.