<p>Applying promptable segmentation foundation models, such as the segment anything model (SAM), to medical image segmentation poses challenges, particularly in addressing irregular morphologies, indistinct boundaries, and sparsely distributed lesions, such as those associated with myopic maculopathy (MM). In this paper, we propose TP-SA3M, a novel two-step framework that customizes SAM for text-prompted MM lesion segmentation. Our approach incorporates explainable medical prior knowledge to guide and enhance the segmentation process. Specifically, we leverage the text–image alignment capabilities of vision-language models to generate these priors. The multi-scale visual–text alignment module learns both local and global semantics of lesions by aligning visual and textual features at different levels. A parameter-efficient prior-injected adapter is incorporated into the SAM encoder, facilitating cross-modal knowledge sharing. These priors also guide the generation of prior-guided prompts (bounding box and points), further refining the segmentation. Additionally, a decoupled mask decoder is employed to mitigate cross-domain degradation caused by domain shifts. Experiments on two public MM datasets demonstrate that our framework achieves state-of-the-art performance.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

TP-SA3M: text prompts-assisted SAM for myopic maculopathy segmentation

  • Tingyao Li,
  • Zehua Jiang,
  • Yixiao Jin,
  • Chunxing Liu,
  • Xiangning Wang,
  • Tingli Chen

摘要

Applying promptable segmentation foundation models, such as the segment anything model (SAM), to medical image segmentation poses challenges, particularly in addressing irregular morphologies, indistinct boundaries, and sparsely distributed lesions, such as those associated with myopic maculopathy (MM). In this paper, we propose TP-SA3M, a novel two-step framework that customizes SAM for text-prompted MM lesion segmentation. Our approach incorporates explainable medical prior knowledge to guide and enhance the segmentation process. Specifically, we leverage the text–image alignment capabilities of vision-language models to generate these priors. The multi-scale visual–text alignment module learns both local and global semantics of lesions by aligning visual and textual features at different levels. A parameter-efficient prior-injected adapter is incorporated into the SAM encoder, facilitating cross-modal knowledge sharing. These priors also guide the generation of prior-guided prompts (bounding box and points), further refining the segmentation. Additionally, a decoupled mask decoder is employed to mitigate cross-domain degradation caused by domain shifts. Experiments on two public MM datasets demonstrate that our framework achieves state-of-the-art performance.