There are massive amounts of materials available on the Internet these days, which creates a huge problem in accurately finding the data. This is because all of the data does not meet the user’s requirements. In most circumstances, we also wish to get the best response to our search query from a variety of online document links. Users may get useful information from a vast quantity of data on the Internet via information retrieval. A search engine is a practical application of information retrieval methods that aids in the discovery of relevant material for the input query. The ranking module of the search engine ranks web pages in order to produce the best results. One cannot stress the significance of ranking in terms of giving users accurate and ideal search results. A number of document similarity metrics, including cosine-similarity, Jaccard-similarity, dice-coefficient, Pearson correlation, and Euclidean distance, are compared in this study. Based on the traits that these documents have in common, these similarity metrics are used to group documents into clusters. The submitted documents’ term-frequency and normalized term-frequency are utilized to create similarity measures in accordance with the formulas outlined by different similarity measure approaches. The F-measure, one of cluster’s performance assessment criteria, is used to compare the effectiveness of various similarity measure algorithms.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Domain-Specific Automatic Cluster Generation

  • Purwa Chaudhary,
  • Mukesh Rawat

摘要

There are massive amounts of materials available on the Internet these days, which creates a huge problem in accurately finding the data. This is because all of the data does not meet the user’s requirements. In most circumstances, we also wish to get the best response to our search query from a variety of online document links. Users may get useful information from a vast quantity of data on the Internet via information retrieval. A search engine is a practical application of information retrieval methods that aids in the discovery of relevant material for the input query. The ranking module of the search engine ranks web pages in order to produce the best results. One cannot stress the significance of ranking in terms of giving users accurate and ideal search results. A number of document similarity metrics, including cosine-similarity, Jaccard-similarity, dice-coefficient, Pearson correlation, and Euclidean distance, are compared in this study. Based on the traits that these documents have in common, these similarity metrics are used to group documents into clusters. The submitted documents’ term-frequency and normalized term-frequency are utilized to create similarity measures in accordance with the formulas outlined by different similarity measure approaches. The F-measure, one of cluster’s performance assessment criteria, is used to compare the effectiveness of various similarity measure algorithms.