Background <p>The internet is a primary source of health information for breast cancer patients, but online content quality varies widely. This study aimed to evaluate the capability of large language models (LLMs), including ChatGPT and Claude, to assess the quality of online Japanese breast cancer treatment information by calculating and comparing their DISCERN scores with those of expert raters.</p> Methods <p>We analyzed 60 Japanese web pages on breast cancer treatments (surgery, chemotherapy, immunotherapy) using the DISCERN instrument. Each page was evaluated by the LLMs ChatGPT and Claude, along with two expert raters. We assessed LLMs evaluation consistency, correlations between LLMs and expert assessments, and relationships between DISCERN scores, Google search rankings, and content length.</p> Results <p>Evaluations by LLMs showed high consistency and moderate to strong correlations with expert assessments (ChatGPT vs Expert: <i>r</i> = 0.65; Claude vs Expert: <i>r</i> = 0.68). LLMs assigned slightly higher scores than expert raters. Chemotherapy pages received the highest quality scores, followed by surgery and immunotherapy. We found a weak negative correlation between Google search ranking and DISCERN scores, and a moderate positive correlation (<i>r</i> = 0.45) between content length and quality ratings.</p> Conclusions <p>This study demonstrates the potential of LLM-assisted evaluation in assessing online health information quality, while highlighting the importance of human expertise. LLMs could efficiently process large volumes of health information but should complement human insight for comprehensive assessments. These findings have implications for improving the accessibility and reliability of breast cancer treatment information.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Assessing the quality of Japanese online breast cancer treatment information using large language models: a comparison of ChatGPT, Claude, and expert evaluations

  • Atsushi Fushimi,
  • Mitsuo Terada,
  • Rie Tahara,
  • Yuko Nakazawa,
  • Madoka Iwase,
  • Tomoko Shibayama,
  • Samy Kotti,
  • Nami Yamashita,
  • Asumi Iesato

摘要

Background

The internet is a primary source of health information for breast cancer patients, but online content quality varies widely. This study aimed to evaluate the capability of large language models (LLMs), including ChatGPT and Claude, to assess the quality of online Japanese breast cancer treatment information by calculating and comparing their DISCERN scores with those of expert raters.

Methods

We analyzed 60 Japanese web pages on breast cancer treatments (surgery, chemotherapy, immunotherapy) using the DISCERN instrument. Each page was evaluated by the LLMs ChatGPT and Claude, along with two expert raters. We assessed LLMs evaluation consistency, correlations between LLMs and expert assessments, and relationships between DISCERN scores, Google search rankings, and content length.

Results

Evaluations by LLMs showed high consistency and moderate to strong correlations with expert assessments (ChatGPT vs Expert: r = 0.65; Claude vs Expert: r = 0.68). LLMs assigned slightly higher scores than expert raters. Chemotherapy pages received the highest quality scores, followed by surgery and immunotherapy. We found a weak negative correlation between Google search ranking and DISCERN scores, and a moderate positive correlation (r = 0.45) between content length and quality ratings.

Conclusions

This study demonstrates the potential of LLM-assisted evaluation in assessing online health information quality, while highlighting the importance of human expertise. LLMs could efficiently process large volumes of health information but should complement human insight for comprehensive assessments. These findings have implications for improving the accessibility and reliability of breast cancer treatment information.