Reasoning with large language models in medicine: a systematic review of techniques, challenges and clinical integration
摘要
Large Language Models (LLMs) have emerged as transformative tools in healthcare, demonstrating unprecedented capabilities in medical reasoning tasks that require complex inference, pattern recognition, and decision-making under uncertainty. This comprehensive review examines the current state of LLMs applications in medical reasoning across diverse clinical contexts, including diagnostic reasoning, clinical decision support, medical imaging analysis, drug discovery, and patient management. We systematically analyze the methodological approaches used to adapt and evaluate LLMs, comparing their performance against traditional clinical decision support systems and human clinicians. We further provide a critical comparative analysis of architectural adaptations, fine-tuning techniques, and domain-specific evaluation protocols. Our review encompasses models such as GPT-4, PaLM, Med-PaLM, and BioGPT, highlighting how model design and training paradigms influence reasoning capabilities, generalization, and clinical applicability. We assess their ability to process multimodal data, generate hypotheses, and provide evidence-based recommendations. Distinct adaptation methods such as prompt engineering, few-shot learning, and reinforcement learning with human feedback are examined for their impact on medical accuracy and robustness. We identify key technical challenges including hallucinations, inherited biases, and accountability issues, along with ethical and deployment barriers. Unlike prior reviews, we emphasize open research problems in Artificial Intelligence (AI), including symbolic integration, and context-aware reasoning framing LLMs as computational systems that push the boundaries of interdisciplinary computer science. While LLMs offer promise for enhancing diagnostic accuracy and decision-making, substantial challenges remain. The review concludes by outlining future directions, including hybrid neuro-symbolic models, rigorous evaluation frameworks, and human-in-the-loop systems to ensure safe, transparent, and fair integration into healthcare. Rather than replacing clinicians, LLMs are best positioned as collaborative AI agents advancing the frontiers of intelligent, assistive technologies in medicine.