<p>This paper aims to build a model that can <b>S</b>egment <b>A</b>nything in 3D medical images, driven by medical terminologies as <b>T</b>ext prompts, termed as <b>SAT</b>. Our main contributions are three-fold: (i) We construct the first multimodal knowledge tree on human anatomy, including <b>6502 anatomical terminologies</b>; Then, we build the largest and most comprehensive segmentation dataset for training, collecting over <b>22K 3D scans</b> from <b>72 datasets</b>, across <b>497 classes</b>, with careful standardization on both image and label space; (ii) We propose to inject medical knowledge into a text encoder via contrastive learning and formulate a large-vocabulary segmentation model that can be prompted by medical terminologies in text form. (iii) We train <b>SAT-Nano</b> (110M parameters) and <b>SAT-Pro</b> (447M parameters). <b>SAT-Pro</b> achieves comparable performance to 72 nnU-Nets—the strongest specialist models trained on each dataset (over 2.2B parameters combined)—over 497 categories. Compared with the interactive approach MedSAM, SAT-Pro consistently outperforms across all 7 human body regions with +7.1% average Dice Similarity Coefficient (DSC) improvement, while showing enhanced scalability and robustness. On 2 external (cross-center) datasets, SAT-Pro achieves higher performance than all baselines (+3.7% average DSC), demonstrating superior generalization ability.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Large-vocabulary segmentation for medical images with text prompts

  • Ziheng Zhao,
  • Yao Zhang,
  • Chaoyi Wu,
  • Xiaoman Zhang,
  • Xiao Zhou,
  • Ya Zhang,
  • Yanfeng Wang,
  • Weidi Xie

摘要

This paper aims to build a model that can Segment Anything in 3D medical images, driven by medical terminologies as Text prompts, termed as SAT. Our main contributions are three-fold: (i) We construct the first multimodal knowledge tree on human anatomy, including 6502 anatomical terminologies; Then, we build the largest and most comprehensive segmentation dataset for training, collecting over 22K 3D scans from 72 datasets, across 497 classes, with careful standardization on both image and label space; (ii) We propose to inject medical knowledge into a text encoder via contrastive learning and formulate a large-vocabulary segmentation model that can be prompted by medical terminologies in text form. (iii) We train SAT-Nano (110M parameters) and SAT-Pro (447M parameters). SAT-Pro achieves comparable performance to 72 nnU-Nets—the strongest specialist models trained on each dataset (over 2.2B parameters combined)—over 497 categories. Compared with the interactive approach MedSAM, SAT-Pro consistently outperforms across all 7 human body regions with +7.1% average Dice Similarity Coefficient (DSC) improvement, while showing enhanced scalability and robustness. On 2 external (cross-center) datasets, SAT-Pro achieves higher performance than all baselines (+3.7% average DSC), demonstrating superior generalization ability.