<p>This study evaluated the effectiveness of large language models (LLMs) and vision-language models (VLMs) in gastroenterology. We used board-style multiple-choice questions to assess the performance of both proprietary and open-source LLMs and VLMs—including GPT, Claude, Gemini, Mistral, Llama, Mixtral, Phi, and Qwen, across different interfaces, computing environments, and levels of compression (quantization). Among the proprietary models, o1-preview (82.0%) and Claude3.5-Sonnet (74.0%) had the highest accuracy, outperforming the top open-source models: Llama3.3-70b (65.7%) and Qwen-2.5-72b (61.0%). Among the small quantized open-source models, the 8-bit Llama 3.2-11b (51.7%) and 6-bit Phi3-14b (48.7%) performed the best, with scores comparable to their full-precision counterparts. Notably, VLM accuracy on image-containing questions improved (~10%) when given human-generated captions, remained unchanged with original images, and declined with LLM-generated captions. Further research is warranted to evaluate model capabilities in real-world clinical decision-making scenarios.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Benchmarking proprietary and open-source language and vision-language models for gastroenterology clinical reasoning

  • Seyed Amir Ahmad Safavi-Naini,
  • Shuhaib Ali,
  • Omer Shahab,
  • Zahra Shahhoseini,
  • Thomas Savage,
  • Sara Rafiee,
  • Jamil S. Samaan,
  • Reem Al Shabeeb,
  • Farah Ladak,
  • Jamie O. Yang,
  • Juan Echavarria,
  • Sumbal Babar,
  • Aasma Shaukat,
  • Samuel Margolis,
  • Nicholas P. Tatonetti,
  • Girish Nadkarni,
  • Bara El Kurdi,
  • Ali Soroush

摘要

This study evaluated the effectiveness of large language models (LLMs) and vision-language models (VLMs) in gastroenterology. We used board-style multiple-choice questions to assess the performance of both proprietary and open-source LLMs and VLMs—including GPT, Claude, Gemini, Mistral, Llama, Mixtral, Phi, and Qwen, across different interfaces, computing environments, and levels of compression (quantization). Among the proprietary models, o1-preview (82.0%) and Claude3.5-Sonnet (74.0%) had the highest accuracy, outperforming the top open-source models: Llama3.3-70b (65.7%) and Qwen-2.5-72b (61.0%). Among the small quantized open-source models, the 8-bit Llama 3.2-11b (51.7%) and 6-bit Phi3-14b (48.7%) performed the best, with scores comparable to their full-precision counterparts. Notably, VLM accuracy on image-containing questions improved (~10%) when given human-generated captions, remained unchanged with original images, and declined with LLM-generated captions. Further research is warranted to evaluate model capabilities in real-world clinical decision-making scenarios.