Purpose <p>This study analyzed responses and readability of generative artificial intelligence (AI) models to questions and recommendations from the 2014 Journal of Neurosurgery: Spine (JNS) guidelines for fusion procedures in the treatment of degenerative lumbar spine disease.</p> Methods <p>Twenty-four questions were generated from JNS guidelines and asked to ChatGPT 4o, Perplexity, Microsoft Copilot, and Gemini. Answers were “concordant” if the response highlighted all points from the JNS guidelines; otherwise, answers were considered “non-concordant” and further sub-categorized as either “insufficient” or “overconclusive.” Responses were evaluated for readability via the Flesch–Kincaid Grade Level, Gunning Fog Index, Simple Measure of Gobbledygook (SMOG) Index, and Flesch Reading Ease test.</p> Results <p>ChatGPT 4o had the highest concordance rate at 66.67%, with non-concordant responses distributed at 16.67% for both insufficient and over-conclusive classifications. Perplexity displayed a 58.33% concordance rate, with 25% insufficient and 16.67% over-conclusive responses. Copilot showed 50% concordance, with 37.5% over-conclusive and 16.67% insufficient responses. Gemini demonstrated 54.17% concordance, with 20.83% insufficient and 25% over-conclusive responses. The Flesch–Kincaid Grade Level scores ranged from 14.03 (Copilot) to 15.66 (Perplexity). The Gunning Fog Index scores varied between 15.15 (Copilot) and 18.13 (Perplexity). The SMOG Index scores ranged from 14.69 (Copilot) to 16.49 (Perplexity). The Flesch Reading Ease scores were low across all models, with Copilot showing the highest score of 20.71.</p> Conclusions <p>ChatGPT 4.0 emerged as the best-performing model in terms of concordance, while Perplexity displayed the highest complexity in text readability. AI can be a valuable adjunct in clinical decision-making but cannot replace clinician judgment.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

The role of generative artificial intelligence in deciding fusion treatment of lumbar degeneration: a comparative analysis and narrative review

  • Taha M. Taka,
  • Christopher E. Collins,
  • Andrew Miner,
  • Isaac Overfield,
  • David Shin,
  • Lauren Seo,
  • Olumide Danisa

摘要

Purpose

This study analyzed responses and readability of generative artificial intelligence (AI) models to questions and recommendations from the 2014 Journal of Neurosurgery: Spine (JNS) guidelines for fusion procedures in the treatment of degenerative lumbar spine disease.

Methods

Twenty-four questions were generated from JNS guidelines and asked to ChatGPT 4o, Perplexity, Microsoft Copilot, and Gemini. Answers were “concordant” if the response highlighted all points from the JNS guidelines; otherwise, answers were considered “non-concordant” and further sub-categorized as either “insufficient” or “overconclusive.” Responses were evaluated for readability via the Flesch–Kincaid Grade Level, Gunning Fog Index, Simple Measure of Gobbledygook (SMOG) Index, and Flesch Reading Ease test.

Results

ChatGPT 4o had the highest concordance rate at 66.67%, with non-concordant responses distributed at 16.67% for both insufficient and over-conclusive classifications. Perplexity displayed a 58.33% concordance rate, with 25% insufficient and 16.67% over-conclusive responses. Copilot showed 50% concordance, with 37.5% over-conclusive and 16.67% insufficient responses. Gemini demonstrated 54.17% concordance, with 20.83% insufficient and 25% over-conclusive responses. The Flesch–Kincaid Grade Level scores ranged from 14.03 (Copilot) to 15.66 (Perplexity). The Gunning Fog Index scores varied between 15.15 (Copilot) and 18.13 (Perplexity). The SMOG Index scores ranged from 14.69 (Copilot) to 16.49 (Perplexity). The Flesch Reading Ease scores were low across all models, with Copilot showing the highest score of 20.71.

Conclusions

ChatGPT 4.0 emerged as the best-performing model in terms of concordance, while Perplexity displayed the highest complexity in text readability. AI can be a valuable adjunct in clinical decision-making but cannot replace clinician judgment.