This study investigates the performance of general and medical-specific Large Language Models (LLMs) in obstetrics and gynecology, focusing on their ability to accurately handle medical multiple-choice questions (MCQs). We evaluated models like Llama2, Mistral, PMC_LLaMA, and BioMistral, to assess and enhance their reliability and accuracy. Despite the expectations, general-purpose models occasionally outperformed specialized medical models. Our methods, including Structural Influence Testing and Contextual Enhancement Testing, demonstrated significant potential in improving model accuracy and reducing misinformation. Specifically, Structural Influence Testing increased Mistral’s accuracy from 40% to 46% and Llama2’s from 28% to 43% with five shots. Contextual Enhancement Testing yielded a 4% accuracy gain for Mistral and 6% for Llama2 using search terms. This research highlights the importance of optimizing LLMs to empower healthcare professionals with precise and reliable medical information, ultimately improving patient outcomes and supporting informed clinical decisions.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Evaluating Large Language Models for Healthcare: Insights from MCQ Evaluation

  • Shuangshuang Lin,
  • Hamzah Bin Osop,
  • Miao Zhang,
  • Xinxian Huang

摘要

This study investigates the performance of general and medical-specific Large Language Models (LLMs) in obstetrics and gynecology, focusing on their ability to accurately handle medical multiple-choice questions (MCQs). We evaluated models like Llama2, Mistral, PMC_LLaMA, and BioMistral, to assess and enhance their reliability and accuracy. Despite the expectations, general-purpose models occasionally outperformed specialized medical models. Our methods, including Structural Influence Testing and Contextual Enhancement Testing, demonstrated significant potential in improving model accuracy and reducing misinformation. Specifically, Structural Influence Testing increased Mistral’s accuracy from 40% to 46% and Llama2’s from 28% to 43% with five shots. Contextual Enhancement Testing yielded a 4% accuracy gain for Mistral and 6% for Llama2 using search terms. This research highlights the importance of optimizing LLMs to empower healthcare professionals with precise and reliable medical information, ultimately improving patient outcomes and supporting informed clinical decisions.