Recommending Influential Authors Using Content-Based Filtering and Network Similarity-A Case Study on Disease-Related Research
摘要
In an era of rapid academic and technological advancement, this research work addresses the challenge of identifying experts from the world’s top author list from different fields of research. In particular, Stanford’s top 2% list of authors has been used for the analysis. Specifically, a recommendation system is proposed that goes beyond textual similarity between user-specific keywords and publication titles by considering various publication-related statistics such as total publications, citations, etc. This study mainly focuses on a specific set of keywords related to cancer, cardio, and gene and also some arbitrary sets of keywords for case studies. The methodology involves collecting authors’ data from individual Google Scholar profiles, pre-processing the data, constructing a network based on keyword similarity, and finally, performing the multilevel community detection followed by the K-means sub-clustering concerned with the publication-based statistics for recommending influential authors. Moreover, the model accuracy has been evaluated through a comparative analysis between individual authors’ profiles with the domain related to the keywords. It has been observed that the exact matches are 8 out of 10 recommendations for the set of keywords related to cancer in the multilevel community detection technique. Analogously, the accuracy of the author’s recommendation for the next set of keywords related to cardio is 0.6, and the accuracy for recommending authors for the set of keywords related to gene is 0.5.