Visual-language reasoning large language models for primary care: advancing clinical decision support through multimodal AI
摘要
This comprehensive review critically evaluates the integration of visual-language reasoning large language models (VL-LLMs) into primary care, with a focus on their transformative potential in medical image analysis, clinical report interpretation, and multimodal decision support systems. By synthesizing advances foundational models such as CLIP, FLAVA, and BLIP with clinical datasets spanning diabetic retinopathy, pulmonary CT analysis, and volumetric segmentation, this work systematically examines how VL-LLMs can bridge the gap between visual diagnostics and natural language processing in clinical workflows. The key challenges identified include mitigation of data biases, improved model explainability, and practical barriers to real-world deployment. The review proposes actionable recommendations for future research, advocating for interdisciplinary collaboration, the establishment of standardized evaluation frameworks, and the prioritization of ethical AI development to ensure equitable clinical translation.