Dual U-shaped cross-modal fusion network for lung infection region segmentation
摘要
For medical image segmentation, there are two major obstacles in constructing high-quality datasets which are the difficulty of acquiring available medical images and financial burden of data annotation. Therefore, we leverage medical text data to compensate for the defects of existing datasets. In this work, a dual U-shaped network based on convolutional neural network and vision transformer is designed to sufficiently achieve the cross-modal feature fusion of image and text. Specifically, one U-shaped branch mainly extracts global features of images and generates the final prediction results. The other one is responsible for processing text information and integrating images with text information. Additionally, we utilize two fusion modules to equip the skip connections and resolve the semantic gaps. Comprehensive experiments have been conducted on two lung image datasets. On QaTa-COV19 dataset, compared to the suboptimal results, our method improves Dice, mIoU by 3.23%, 3.89% and reduces ASD, HD95 by 2.73mm, 10.76mm.