Z Visual-language foundation models for medical and clinical diagnosis and treatments
摘要
This comprehensive review explores the transformative potential of visual-language foundation models (VL-FMs) in medical and clinical applications. By integrating deep visual perception with natural language understanding, these models enable context-aware interpretation of radiological, pathological, and ophthalmic images alongside clinical text. Technical advancements are analyzed, including transformer architectures, multi-modal fusion strategies, and knowledge integration frameworks. Case studies illustrate applications in disease diagnosis (e.g., diabetic retinopathy, lung cancer), image segmentation (e.g., brain tumors, vascular structures), and clinical decision support. Existing challenges such as data privacy, interpretability, and computational efficiency are discussed, alongside future directions for developing explainable and generalizable AI systems in precision medicine.