This work addresses the challenge of navigating the vast biomedical literature, particularly in the context of Sudan. With 1,575 articles about Sudan in PubMed alone, traditional list-based search engines fail to provide effective knowledge discovery. The project, led by Professor Ahmed Alsafi, set out the task of analyzing critical information contained in more than 19,000 Sudanese biomedical documents and health reports. To manage this extensive corpus, the project uses advanced techniques such as content analysis, natural language processing (NLP), topic modeling, and correlation analysis. Initially, non-English material is extracted, eliminated, and then duplicates are removed, reducing the dataset to approximately 13,000 documents. Latent Dirichlet allocation (LDA) is then used to generate thematic models, and the papers are clustered into 40 significant research themes. Accordingly, NetworkX is used to create co-author networks and topic networks. These networks help identify similar papers and establish profiles of researchers, providing a deeper understanding of their research interests and careers. Specifically, this work exemplifies the power of comprehensive data mining techniques to extract valuable insights from large collections of biological research papers. It found out that the most common subjects in Sudan are the exposure of the Sudanese population to pulmonary tuberculosis, variations in milk composition among Sudanese provinces, and the prevalence of mental health issues among refugee populations, as they have the highest number of documents. It also discovered that there are 46,545 authors in the dataset. This impact transcends the boundaries of the Sudanese research landscape and fosters collaboration and knowledge sharing within the broader biomedical research community.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Cognitive Network Analysis and Topic Modeling for Bibliography Data of Biomedical Literature of Sudan

  • Ruba Salaheldin Osman,
  • Reham Jamal Taha,
  • Marwan Adam,
  • Sharief Babiker

摘要

This work addresses the challenge of navigating the vast biomedical literature, particularly in the context of Sudan. With 1,575 articles about Sudan in PubMed alone, traditional list-based search engines fail to provide effective knowledge discovery. The project, led by Professor Ahmed Alsafi, set out the task of analyzing critical information contained in more than 19,000 Sudanese biomedical documents and health reports. To manage this extensive corpus, the project uses advanced techniques such as content analysis, natural language processing (NLP), topic modeling, and correlation analysis. Initially, non-English material is extracted, eliminated, and then duplicates are removed, reducing the dataset to approximately 13,000 documents. Latent Dirichlet allocation (LDA) is then used to generate thematic models, and the papers are clustered into 40 significant research themes. Accordingly, NetworkX is used to create co-author networks and topic networks. These networks help identify similar papers and establish profiles of researchers, providing a deeper understanding of their research interests and careers. Specifically, this work exemplifies the power of comprehensive data mining techniques to extract valuable insights from large collections of biological research papers. It found out that the most common subjects in Sudan are the exposure of the Sudanese population to pulmonary tuberculosis, variations in milk composition among Sudanese provinces, and the prevalence of mental health issues among refugee populations, as they have the highest number of documents. It also discovered that there are 46,545 authors in the dataset. This impact transcends the boundaries of the Sudanese research landscape and fosters collaboration and knowledge sharing within the broader biomedical research community.