Development of a Strategy for Duplication Search Based on Multiple Hierarchically Organized Approaches
摘要
Data deduplication is essential for records in business systems, as duplicate entries can distort analyses, harm relationships, and generate unnecessary costs. As organizations increasingly rely on diverse and large-scale data sources, the challenge of identifying and consolidating duplicate records has grown, making efficient deduplication a critical factor for maintaining data quality and operational efficiency. This paper proposes a hierarchical data deduplication approach to address this industry-wide challenge. The proposed approach, referred to here as the engine, integrates specialized rules, similarity measures, and large language models (LLMs) in sequential layers to analyze data and identify duplications accurately and efficiently. By leveraging a structured combination of techniques, the approach overcomes the limitations of traditional methods, which often depend on isolated and less adaptable solutions. The engine balances performance and accuracy, handling both straightforward and complex cases of duplication. Experimental results using a customer-simulated database demonstrate that integrating these advanced techniques improves adaptability and precision compared to conventional methods. This strategy shows significant potential for large-scale data deduplication, offering a more robust solution to an industry-wide challenge.