Extracting relevant treatment strategies for COVID-19 from biomedical literature is crucial for efficient knowledge discovery. This study proposes a text mining approach to identify key treatment modalities from PubMed abstracts using one-cluster clustering and term frequency (tf) representation. The methodology involves sentence segmentation, text preprocessing, feature selection using Chi-Square (χ2), and constraint-based k-means clustering. The experimental results demonstrate that tf representation outperforms binary representation, achieving a recall score of 0.795 vs. 0.750, indicating improved identification of treatment-related insights. By employing constraint-based k-means clustering with k = 1, the model effectively consolidates all relevant information into a single cluster, preserving essential treatment details. The integration of Chi-Square feature selection further refines the extracted data by identifying the most significant terms. This approach provides a systematic and scalable method for biomedical text analysis.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Leveraging PubMed Abstracts for Identifying COVID-19 Treatment Modalities

  • Pornpavit Donsena,
  • Jantima Polpinij,
  • Bancha Luaphol

摘要

Extracting relevant treatment strategies for COVID-19 from biomedical literature is crucial for efficient knowledge discovery. This study proposes a text mining approach to identify key treatment modalities from PubMed abstracts using one-cluster clustering and term frequency (tf) representation. The methodology involves sentence segmentation, text preprocessing, feature selection using Chi-Square (χ2), and constraint-based k-means clustering. The experimental results demonstrate that tf representation outperforms binary representation, achieving a recall score of 0.795 vs. 0.750, indicating improved identification of treatment-related insights. By employing constraint-based k-means clustering with k = 1, the model effectively consolidates all relevant information into a single cluster, preserving essential treatment details. The integration of Chi-Square feature selection further refines the extracted data by identifying the most significant terms. This approach provides a systematic and scalable method for biomedical text analysis.