Link Discovery plays a vital role in enhancing connections between data sources within the Linked Open Data Cloud. A key aspect of this process is Entity Resolution (ER), which focuses on identifying owl:sameAs relationships between different entity descriptions representing the same real-world object. Recent works on ER have investigated the use of Large Language Models (LLMs) for Entity Matching, showing promising results. In the simplest case, a pair of entities is given as input to an LLM, asking whether these entities match or not. A recent approach introduced SELECT prompts, which include the query entity along with multiple candidates generated by a Blocking method. However, this increases the complexity of the questions posed to LLMs, while being susceptible to position bias among the presented candidates. To address these issues, we introduce AvengER, a novel approach to SELECT prompts for LLM-based Matching that effectively handles both unsupervised and supervised settings. For the former, we introduce a hybrid approach that utilizes an ensemble of open-source medium-size models (8b) and also selectively leverages (in just 12% of the cases) an external larger-size Judge (32b), thus balancing high accuracy with computational efficiency. For the latter, we use a dataset containing data from multiple domains to fine-tune a medium-size model so that it surpasses both its pre-trained version and pre-trained large-size models (GPT-3.5).

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

AvengER: Ensembling and Fine-Tuning LLMs for SELECT Prompts in Entity Resolution

  • Alexandros Zeakis,
  • George Papadakis,
  • Dimitrios Skoutas,
  • Manolis Koubarakis

摘要

Link Discovery plays a vital role in enhancing connections between data sources within the Linked Open Data Cloud. A key aspect of this process is Entity Resolution (ER), which focuses on identifying owl:sameAs relationships between different entity descriptions representing the same real-world object. Recent works on ER have investigated the use of Large Language Models (LLMs) for Entity Matching, showing promising results. In the simplest case, a pair of entities is given as input to an LLM, asking whether these entities match or not. A recent approach introduced SELECT prompts, which include the query entity along with multiple candidates generated by a Blocking method. However, this increases the complexity of the questions posed to LLMs, while being susceptible to position bias among the presented candidates. To address these issues, we introduce AvengER, a novel approach to SELECT prompts for LLM-based Matching that effectively handles both unsupervised and supervised settings. For the former, we introduce a hybrid approach that utilizes an ensemble of open-source medium-size models (8b) and also selectively leverages (in just 12% of the cases) an external larger-size Judge (32b), thus balancing high accuracy with computational efficiency. For the latter, we use a dataset containing data from multiple domains to fine-tune a medium-size model so that it surpasses both its pre-trained version and pre-trained large-size models (GPT-3.5).