Anatomy-Aware Mixture of Experts for Medical Vision-Language Pre-training
摘要
Medical Vision-Language Pre-training methods effectively leverage supervisory information from reports text to enhance representation learning. However, many existing approaches depend on a static vision encoder and mainly focus on designing specific optimization objectives for inter-modality and intra-modality learning. The parameters of the vision encoder remain unchanged during the pre-training, fine-tuning, and inference, which limits the flexibility and performance of the model across different downstream tasks. In order to overcome these limitations, we first introduce a three-stage sparse vision encoder, specifically designed to extract visual representations, and propose a novel Anatomy-aware Mixture of Experts (AnaMoE) framework that enables dynamic expansion of the neural network’s capacity without increasing computational costs. AnaMoE integrates a sparse mixture of experts with a threshold-based adaptation mechanism and decouples anatomical features through a collaborative expert fusion module. Extensive experimental evaluations across five medical image datasets and four distinct tasks—temporal image classification, medical image semantic segmentation, linear image classification, and zero-shot image classification—show that AnaMoE greatly surpasses existing methods, demonstrating its superior flexibility and effectiveness in medical vision-language pre-training.