Medical Segment Anything Model (MedSAM): methodological innovations, evaluation practices, and clinical applications
摘要
The Medical Segment Anything Model (MedSAM) has emerged as a prominent foundation model for medical image segmentation, prompting rapid growth in related research. Despite this proliferation, the literature remains fragmented across multimodal paradigms and lacks rigorous synthesis of methodological innovations, clinical safety evidence, algorithmic bias, and real-world implementation barriers—gaps that limit the field’s ability to translate algorithmic progress into clinical impact.
ObjectiveTo systematically identify, classify, and synthesize published research on MedSAM and its derivatives, providing a comprehensive overview of methodological developments, evaluation findings, and clinical translation progress.
MethodsThis scoping review followed the PRISMA-ScR guideline. We conducted a systematic search across six major databases (PubMed, Embase, Web of Science, Scopus, IEEE Xplore, ACM Digital Library) through June 2025.
ResultsFrom 1,482 initial records, 45 studies met inclusion criteria. Innovation-oriented studies (n = 32) were classified along four methodological axes spanning multimodal paradigms: Architecture Enhancement (encoder/decoder refinements for efficiency and precision), Prompt Strategy (textual, spatial, and multimodal prompting mechanisms), Efficiency and Deployability (quantization, distillation, and system integration), and Pipeline and Application (integration into clinical workflows). Evaluation-oriented studies (n = 13) revealed that MedSAM outperforms general-purpose SAM by 5–20% on structured modalities (CT/MRI) but lags behind task-specific networks for delicate structures. Notably, fairness analyses from two studies (combined n = 1,757 subjects) report Dice disparities of up to 10% across sex, age, and BMI strata, raising concerns about potential demographic bias that warrants further investigation in larger, multi-center cohorts before firm clinical-safety conclusions can be drawn. Real-world deployment is further constrained by prompt sensitivity (5–15% Dice variation), cross-modality domain shift, regulatory compliance gaps (FDA/CE marking, HIPAA, GDPR), and absent multi-center prospective validation, all of which represent principal implementation hurdles for clinical adoption.
ConclusionsMedSAM represents a significant advancement in medical image segmentation, demonstrating a clear trajectory from large-scale architectures toward lightweight, prompt-aware, and deployable models. This synthesis provides a structured critical appraisal indicating that the evidence base remains at the proof-of-concept stage, with most studies retrospective and single-center, and emerging signals of algorithmic bias, unresolved cross-modality failure modes, and incomplete regulatory pathways represent safety and implementation considerations that should be addressed before multimodal MedSAM-based systems are more broadly deployed in clinical imaging workflows.