Existing works on rule discovery often yield an overwhelming number of rules with a high computational cost. Moreover, few studies consider the impact of data quality on the reliability of rules produced. This paper studies the top-k rule discovery problem that returns a set of top-k reliable rules. We introduce a novel scoring function that incorporates data quality, providing a more comprehensive and fine-grained assessment of rules. Based on this scoring function, we develop an efficient algorithm for discovering top-k high-quality rules in relational data. To enhance the performance of our algorithm, we propose effective pruning and optimization strategies that significantly reduce the search space and improve computational efficiency; additionally, we parallelize it for further acceleration. Through extensive experiments on real-life datasets, we verify the effectiveness and efficiency of the proposed method, confirming its practical utility in data quality-driven top-k rule discovery.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Data Quality-Driven Top-k Rule Discovery

  • Ziyan Han,
  • Wanjia Chen,
  • Yunpeng Han

摘要

Existing works on rule discovery often yield an overwhelming number of rules with a high computational cost. Moreover, few studies consider the impact of data quality on the reliability of rules produced. This paper studies the top-k rule discovery problem that returns a set of top-k reliable rules. We introduce a novel scoring function that incorporates data quality, providing a more comprehensive and fine-grained assessment of rules. Based on this scoring function, we develop an efficient algorithm for discovering top-k high-quality rules in relational data. To enhance the performance of our algorithm, we propose effective pruning and optimization strategies that significantly reduce the search space and improve computational efficiency; additionally, we parallelize it for further acceleration. Through extensive experiments on real-life datasets, we verify the effectiveness and efficiency of the proposed method, confirming its practical utility in data quality-driven top-k rule discovery.