In this tutorial we present the results of researching, designing, implementing, and deploying data deduplication pipelines for customer records in a big financial institution. The tutorial is based on our experience gained within a R&D project. In the project we developed two deduplication pipelines. The first one is based on statistical modeling, whereas the second one is based on machine learning. Both pipelines were extensively tested on a real data set including customer records. The pipeline based on statistical modeling has already been deployed in the production system of the financial institution and processes batches of over 20 million of customer records.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

On Customer Data Deduplication - Research vs. Industrial Perspective:

  • Witold Andrzejewski,
  • Bartosz Bębel,
  • Paweł Boiński,
  • Robert Wrembel

摘要

In this tutorial we present the results of researching, designing, implementing, and deploying data deduplication pipelines for customer records in a big financial institution. The tutorial is based on our experience gained within a R&D project. In the project we developed two deduplication pipelines. The first one is based on statistical modeling, whereas the second one is based on machine learning. Both pipelines were extensively tested on a real data set including customer records. The pipeline based on statistical modeling has already been deployed in the production system of the financial institution and processes batches of over 20 million of customer records.