Background <p>Recent studies emphasize online health information quality. However, little research focuses on Arabic content for emergency conditions like heart disease, hypertension, and stroke despite the high demand for this information. This study aims to address this gap by using artificial intelligence to create a benchmark dataset for Arabic health information and evaluate its quality.</p> Methods <p>We assessed the quality of health information across three criteria: source quality, treatment quality, and content trustworthiness. The Kruskal-Wallis test was used to analyze quality differences across content types (General Health Information, Medical Advice, Treatment Description) and website categories (Government, Journalistic, Portal, Professional). Data augmentation techniques such as paraphrasing, back translation, and RandAugment were also employed to enhance quality assessment using the Arabic BERT model. The study also proposes a novel architecture termed the Mixture of Classification. In this approach, each health document is processed in parallel by three instances of an Arabic BERT model: the first identifies the type of health information, the second determines the category of the provider, and the third estimates a continuous quality score via regression.</p> Results <p>Significant quality differences were observed among website categories and content types. Portal sites achieved the highest mean score (11.64), while Journalistic sites scored the lowest (3.33). Treatment Descriptions had the highest score (18.74), while Medical Advice scored the lowest (4.73). These differences were statistically significant with large effect sizes (Cohen’s <InlineEquation ID="IEq1"> <EquationSource Format="TEX">\(f = 0.409\)</EquationSource> </InlineEquation> for categories; <InlineEquation ID="IEq2"> <EquationSource Format="TEX">\(f = 0.288\)</EquationSource> </InlineEquation> for content types), indicating substantial practical impact. Paraphrasing showed the best performance, with 89% accuracy, 85% F1 score, 76% precision, and 82% recall with high confidence and minimal class imbalance. A notable 36.3% accuracy gap was identified between low-quality (95.01%) and high-quality (58.71%) content using RandAugment data. These findings highlight the importance of content type, provider category, and quality score for improving health information search rankings.</p> Conclusion <p>Content type, provider category, and quality score are key factors in enhancing the ranking of Arabic health information. Paraphrased data augmentation contributes to improved model reliability in distinguishing between quality classes. Future research should extend this approach to other languages and health-related topics. However, a major challenge remains: achieving a balanced dataset, particularly for binary classification between high- and low-quality content, as well as for the new classification based on provider category, content type, and quality score. The goal is to ensure an equal distribution across all these categories.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Artificial intelligence for automating the establishment of an Arabic benchmark dataset for enhancing health information quality assessment

  • Yousef Baqraf,
  • Pantea Keikhosrokiani,
  • Yu-N. Cheah,
  • Amany Alomarji,
  • Fatima Baqraf

摘要

Background

Recent studies emphasize online health information quality. However, little research focuses on Arabic content for emergency conditions like heart disease, hypertension, and stroke despite the high demand for this information. This study aims to address this gap by using artificial intelligence to create a benchmark dataset for Arabic health information and evaluate its quality.

Methods

We assessed the quality of health information across three criteria: source quality, treatment quality, and content trustworthiness. The Kruskal-Wallis test was used to analyze quality differences across content types (General Health Information, Medical Advice, Treatment Description) and website categories (Government, Journalistic, Portal, Professional). Data augmentation techniques such as paraphrasing, back translation, and RandAugment were also employed to enhance quality assessment using the Arabic BERT model. The study also proposes a novel architecture termed the Mixture of Classification. In this approach, each health document is processed in parallel by three instances of an Arabic BERT model: the first identifies the type of health information, the second determines the category of the provider, and the third estimates a continuous quality score via regression.

Results

Significant quality differences were observed among website categories and content types. Portal sites achieved the highest mean score (11.64), while Journalistic sites scored the lowest (3.33). Treatment Descriptions had the highest score (18.74), while Medical Advice scored the lowest (4.73). These differences were statistically significant with large effect sizes (Cohen’s \(f = 0.409\) for categories; \(f = 0.288\) for content types), indicating substantial practical impact. Paraphrasing showed the best performance, with 89% accuracy, 85% F1 score, 76% precision, and 82% recall with high confidence and minimal class imbalance. A notable 36.3% accuracy gap was identified between low-quality (95.01%) and high-quality (58.71%) content using RandAugment data. These findings highlight the importance of content type, provider category, and quality score for improving health information search rankings.

Conclusion

Content type, provider category, and quality score are key factors in enhancing the ranking of Arabic health information. Paraphrased data augmentation contributes to improved model reliability in distinguishing between quality classes. Future research should extend this approach to other languages and health-related topics. However, a major challenge remains: achieving a balanced dataset, particularly for binary classification between high- and low-quality content, as well as for the new classification based on provider category, content type, and quality score. The goal is to ensure an equal distribution across all these categories.