Many modern Entity Resolution (ER) systems leverage metadata about the reference data to facilitate processing and making equivalence decisions. This historically has required that each source of input data be pre-processed individually to conform to a common metadata alignment, have data cleansing applied, and the data condensed into a singular dataset to be submitted to an ER process. These are costly processes and require additional passes of the input data prior to ER decisioning. This paper expands on the concept of context within ER to replace the need for metadata alignment and preprocessing that was introduced in previous literature. Leveraging the four (4) points within ER processes for which context can be extracted, and introducing a new n-gram similarity method allows for better more intelligent corrections to be performed to the data during processing without the need for preprocessing or metadata. This paper defines tested methods to enhance the initial context extraction to reduce erroneous automated corrections that thusly impact the final clustering results of the ER system and provides empirical results to support the effectiveness of these methods.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Using Linkage Context for Automated Correction in Unsupervised Entity Resolution

  • Fumiko Kobayashi,
  • John R. Talburt

摘要

Many modern Entity Resolution (ER) systems leverage metadata about the reference data to facilitate processing and making equivalence decisions. This historically has required that each source of input data be pre-processed individually to conform to a common metadata alignment, have data cleansing applied, and the data condensed into a singular dataset to be submitted to an ER process. These are costly processes and require additional passes of the input data prior to ER decisioning. This paper expands on the concept of context within ER to replace the need for metadata alignment and preprocessing that was introduced in previous literature. Leveraging the four (4) points within ER processes for which context can be extracted, and introducing a new n-gram similarity method allows for better more intelligent corrections to be performed to the data during processing without the need for preprocessing or metadata. This paper defines tested methods to enhance the initial context extraction to reduce erroneous automated corrections that thusly impact the final clustering results of the ER system and provides empirical results to support the effectiveness of these methods.