Background <p>The Bosniak classification system is widely used to assess malignancy risk in renal cystic lesions, yet inter-observer variability poses significant challenges. Large language models (LLMs) may offer a&#xa0;standardized approach to classification when provided with textual descriptions, such as those found in radiology reports.</p> Objective <p>This study evaluated the performance of five LLMs—GPT‑4 (ChatGPT), Gemini, Copilot, Perplexity, and NotebookLM—in classifying renal cysts based on synthetic textual descriptions mimicking CT report content.</p> Methods <p>A&#xa0;synthetic dataset of 100 diagnostic scenarios (20&#xa0;cases per Bosniak category) was constructed using established radiological criteria. Each LLM was evaluated using zero-shot and few-shot prompting strategies, while NotebookLM employed retrieval-augmented generation (RAG). Performance metrics included accuracy, sensitivity, and specificity. Statistical significance was assessed using McNemar’s and chi-squared tests.</p> Results <p>GPT‑4 achieved the highest accuracy (87% zero-shot, 99% few-shot), followed by Copilot (81–86%), Gemini (55–69%), and Perplexity (43–69%). NotebookLM, tested only under RAG conditions, reached 87% accuracy. Few-shot learning significantly improved performance (<i>p</i> &lt; 0.05). Classification of Bosniak IIF lesions remained challenging across models.</p> Conclusion <p>When provided with well-structured textual descriptions, LLMs can accurately classify renal cysts. Few-shot prompting significantly enhances performance. However, persistent difficulties in classifying borderline lesions such as Bosniak IIF highlight the need for further refinement and real-world validation.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Bosniak classification of renal cysts using large language models: a comparative study

  • Ibrahim Hacibey,
  • Esat Kaba

摘要

Background

The Bosniak classification system is widely used to assess malignancy risk in renal cystic lesions, yet inter-observer variability poses significant challenges. Large language models (LLMs) may offer a standardized approach to classification when provided with textual descriptions, such as those found in radiology reports.

Objective

This study evaluated the performance of five LLMs—GPT‑4 (ChatGPT), Gemini, Copilot, Perplexity, and NotebookLM—in classifying renal cysts based on synthetic textual descriptions mimicking CT report content.

Methods

A synthetic dataset of 100 diagnostic scenarios (20 cases per Bosniak category) was constructed using established radiological criteria. Each LLM was evaluated using zero-shot and few-shot prompting strategies, while NotebookLM employed retrieval-augmented generation (RAG). Performance metrics included accuracy, sensitivity, and specificity. Statistical significance was assessed using McNemar’s and chi-squared tests.

Results

GPT‑4 achieved the highest accuracy (87% zero-shot, 99% few-shot), followed by Copilot (81–86%), Gemini (55–69%), and Perplexity (43–69%). NotebookLM, tested only under RAG conditions, reached 87% accuracy. Few-shot learning significantly improved performance (p < 0.05). Classification of Bosniak IIF lesions remained challenging across models.

Conclusion

When provided with well-structured textual descriptions, LLMs can accurately classify renal cysts. Few-shot prompting significantly enhances performance. However, persistent difficulties in classifying borderline lesions such as Bosniak IIF highlight the need for further refinement and real-world validation.