Entity alignment (EA) is a crucial process in integrating data from multiple sources, facilitating Knowledge Discovery (KD). Despite advances in EA techniques, selecting the appropriate algorithm for downstream KD tasks remains challenging due to several issues. These issues include domain entities alignment difficulties, the impact on KD tasks, and bias in data distribution. This paper presents a framework to address these challenges by providing a systematic approach to evaluate the impact of different EA algorithms based on three critical aspects: quality of alignment, information retrieved through alignment, and information imbalance or bias introduced through alignment. Our framework enables users to make informed decisions about algorithm selection, ensuring reliable, effective, and balanced KD. We demonstrate the application of the framework using a digital humanities case study, where the KD task involves enriching information about colonial collections. The choice of such a sensitive and historically imbalanced use-case allows us to highlight how the proposed framework helps identify suitable algorithms and to emphasis the importance of understanding the propagated information biases introduced through data alignment.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A Framework for Evaluating Entity Alignment Impact on Downstream Knowledge Discovery

  • Sarah Binta Alam Shoilee,
  • Victor de Boer,
  • Jacco van Ossenbruggen

摘要

Entity alignment (EA) is a crucial process in integrating data from multiple sources, facilitating Knowledge Discovery (KD). Despite advances in EA techniques, selecting the appropriate algorithm for downstream KD tasks remains challenging due to several issues. These issues include domain entities alignment difficulties, the impact on KD tasks, and bias in data distribution. This paper presents a framework to address these challenges by providing a systematic approach to evaluate the impact of different EA algorithms based on three critical aspects: quality of alignment, information retrieved through alignment, and information imbalance or bias introduced through alignment. Our framework enables users to make informed decisions about algorithm selection, ensuring reliable, effective, and balanced KD. We demonstrate the application of the framework using a digital humanities case study, where the KD task involves enriching information about colonial collections. The choice of such a sensitive and historically imbalanced use-case allows us to highlight how the proposed framework helps identify suitable algorithms and to emphasis the importance of understanding the propagated information biases introduced through data alignment.