The “Server Access Pattern Analysis” Based on Different “Weblogs Classification” Methods
摘要
The academic and industrial sectors have shown a significant increase in interest in web usage mining due to its potential for applications. This work offers a thorough taxonomy of the work being done in this subject, including research projects and the way people use the internet today. During the functioning of the system, extensive information and runtime statistics are recorded in the logs. It includes a log message along with a timestamp that indicates what has taken place in the system. A recent review of the work that has previously performed is also presented, considering various data mining methodologies including clustering and classification. The primary focus of this study is the data preprocessing step, which includes activities such as data cleansing techniques and field extraction at the preliminary stage of ‘web usage mining’. ‘The procedure of ‘extracting’ fields from the log file’s single line is carried out by the field extraction algorithm’. The evaluated data is cleaned by an algorithm that removes items that are incorrect or unnecessary.