Almost all data mining algorithms result in better data model representation when working with large amount of good quality data, whereas with small amount of poor quality data result in bad data model. The work in this research tends to improve the mining algorithm's performance. An approach is proposed by using data reduction concepts to get summarized data keeping the same knowledge in the detailed data. The proposed approach includes several phases; In the initial stage, data were collected which are real data representing information of students from different schools in Najaf, the data were converted into a subject-oriented table, the warehouse model design was created to represent the data (galaxy model), the data cube was built, and finally the mining algorithm was applied, analyzed and then implemented. The most important criteria in various data mining algorithms (clustering, classification, and others) are the result accuracy and the speed of the algorithm (time complexity). The results obtained from the proposed approach when applied on the collected data (more than 18,266 records) showed good improvement in both time complexity and accuracy. Both WEKA and C# programming language were used to implement the proposed approach.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A Complete Framework for Improving Data Mining Algorithms

  • Zahraa Al-Barmani

摘要

Almost all data mining algorithms result in better data model representation when working with large amount of good quality data, whereas with small amount of poor quality data result in bad data model. The work in this research tends to improve the mining algorithm's performance. An approach is proposed by using data reduction concepts to get summarized data keeping the same knowledge in the detailed data. The proposed approach includes several phases; In the initial stage, data were collected which are real data representing information of students from different schools in Najaf, the data were converted into a subject-oriented table, the warehouse model design was created to represent the data (galaxy model), the data cube was built, and finally the mining algorithm was applied, analyzed and then implemented. The most important criteria in various data mining algorithms (clustering, classification, and others) are the result accuracy and the speed of the algorithm (time complexity). The results obtained from the proposed approach when applied on the collected data (more than 18,266 records) showed good improvement in both time complexity and accuracy. Both WEKA and C# programming language were used to implement the proposed approach.