Multi-Perspective Text-Guided Multimodal Fusion Network for Brain Tumor Segmentation
摘要
In brain tumor segmentation research, there is considerable interest in fully exploring the potential of all modalities. Most multi-modal fusion segmentation methods rely mainly on traditional discrete label representation learning, focusing solely on utilizing image data for segmentation tasks. The creation of multi-modal medical imaging datasets requires specialized knowledge and time-consuming, making it difficult to achieve large-scale datasets. Networks that rely solely on images are prone to bottlenecks due to the limitations in the quantity and quality of available images. With the emergence of pre-trained visual-language models, the establishment of spatial structural consistency between image and text data enables text information to serve as prompt, guiding models to achieve significant performance. This approach also aids in establishing spatial structural consistency between image and text data. Inspired by these insights, we propose a multi-perspective text-guided multi-modal fusion segmentation network. This network provides semantic guidance for feature extraction fusion and output result deorthogonality through modal and class text prompts, respectively. Our method outperforms existing approaches, achieving superior segmentation performance as demonstrated by evaluation on the BraTS2020 and BraTS2021 datasets.