<p>This exploratory study examined the effectiveness of ChatGPT-4o, Claude 3.5 Sonnet, and ChatGPT-o1 in developing anesthesia plans for critical cases. Personalized anesthesia plans are essential for ensuring surgical safety and patient satisfaction. These artificial intelligence (AI) models can understand and generate anesthesia-related information. The study included a panel of five anesthesia experts, each with over ten years of experience. They qualitatively and quantitatively assessed the capabilities of the three models in formulating anesthesia plans for critical cases. The results showed no significant differences in the response quality, relevance, and applicability scores among the models; however, variations were observed in the error types and severity. ChatGPT-o1 surpassed the other models in terms of content relevance and information accuracy, demonstrating a lower error rate and higher suitability for clinical application. As an initial investigation in this rapidly evolving field, this research provides preliminary insights while acknowledging the need for further validation in clinical settings before implementation.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

An Exploratory Comparison of AI Models for Preoperative Anesthesia Planning: Assessing ChatGPT-4o, Claude 3.5 Sonnet, and ChatGPT-o1 in Clinical Scenario Analysis

  • Bing Wang,
  • Yue Tian,
  • Xue Ting Wang

摘要

This exploratory study examined the effectiveness of ChatGPT-4o, Claude 3.5 Sonnet, and ChatGPT-o1 in developing anesthesia plans for critical cases. Personalized anesthesia plans are essential for ensuring surgical safety and patient satisfaction. These artificial intelligence (AI) models can understand and generate anesthesia-related information. The study included a panel of five anesthesia experts, each with over ten years of experience. They qualitatively and quantitatively assessed the capabilities of the three models in formulating anesthesia plans for critical cases. The results showed no significant differences in the response quality, relevance, and applicability scores among the models; however, variations were observed in the error types and severity. ChatGPT-o1 surpassed the other models in terms of content relevance and information accuracy, demonstrating a lower error rate and higher suitability for clinical application. As an initial investigation in this rapidly evolving field, this research provides preliminary insights while acknowledging the need for further validation in clinical settings before implementation.