Purpose <p>Research on vision-language models (VLMs) in the medical field has recently increased. However, while multifaceted evaluation is necessary to avoid the high risks associated with misdiagnosis, artificial intelligence (AI)-assisted mammogram report generation remains insufficient, with no studies on objective and subjective generation. We aimed to develop an AI system that generates mammogram reports and to verify the impact of differences between objective and subjective evaluation methods on the interpretation of this AI system.</p> Materials and methods <p>We used a public dataset consisting of mammograms and their reports, preparing question prompts and performing low-rank adaptation tuning on Qwen2.5(7B). We analyzed the Breast Imaging Reporting and Data System (BI-RADS) and findings agreement rate, Recall-Oriented Understudy for Gisting Evaluation (ROUGE), and Bilingual Evaluation Understudy (BLEU) for the objective evaluation. A breast clinician performed score-based evaluations of generated reports as subjective assessments. Finally, we analyzed samples of inconsistent objective and subjective results.</p> Results <p>The BI-RADS agreement rate was 58.1%. Findings were accurately included in generated reports at 76.7% for mass and 81.4% for calcification. ROUGE-L F1 and overall BLEU were 0.672 and 0.542, respectively. Although ROUGE-L F1 or overall BLEU was above average, two samples received low scores in the generated report evaluation; these were considered over- and under-estimation.</p> Conclusion <p>We developed an AI model that generates mammogram reports. In assessing this model, we found that multifaceted objective and subjective medical VLM evaluations are necessary for determining over- and under-estimation.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Verification of the impact of differences between objective and subjective evaluation methods on the interpretation of artificial intelligence systems generating mammogram reports using vision-language model

  • Chiharu Kai,
  • Hideaki Tamori,
  • Yuta Hirono,
  • Sachi Ishizuka,
  • Satoshi Kondo,
  • Tsunehiro Ohtsuka,
  • Satoshi Kasai

摘要

Purpose

Research on vision-language models (VLMs) in the medical field has recently increased. However, while multifaceted evaluation is necessary to avoid the high risks associated with misdiagnosis, artificial intelligence (AI)-assisted mammogram report generation remains insufficient, with no studies on objective and subjective generation. We aimed to develop an AI system that generates mammogram reports and to verify the impact of differences between objective and subjective evaluation methods on the interpretation of this AI system.

Materials and methods

We used a public dataset consisting of mammograms and their reports, preparing question prompts and performing low-rank adaptation tuning on Qwen2.5(7B). We analyzed the Breast Imaging Reporting and Data System (BI-RADS) and findings agreement rate, Recall-Oriented Understudy for Gisting Evaluation (ROUGE), and Bilingual Evaluation Understudy (BLEU) for the objective evaluation. A breast clinician performed score-based evaluations of generated reports as subjective assessments. Finally, we analyzed samples of inconsistent objective and subjective results.

Results

The BI-RADS agreement rate was 58.1%. Findings were accurately included in generated reports at 76.7% for mass and 81.4% for calcification. ROUGE-L F1 and overall BLEU were 0.672 and 0.542, respectively. Although ROUGE-L F1 or overall BLEU was above average, two samples received low scores in the generated report evaluation; these were considered over- and under-estimation.

Conclusion

We developed an AI model that generates mammogram reports. In assessing this model, we found that multifaceted objective and subjective medical VLM evaluations are necessary for determining over- and under-estimation.