Enhancing multimodal named entity recognition with multi-granularity knowledge distillation
摘要
Multimodal Named Entity Recognition (MNER) is essential for effective information extraction, yet traditional methods encounter several challenges. These include mismatches between text and images, insufficient utilization of image data, neglect of critical features, and difficulties in aligning semantic levels across modalities. Typically, these approaches focus on aligning and fusing multimodal data without integrating external knowledge. To address these shortcomings, we propose a Multi-Granularity Knowledge Distillation model for MNER (MKD), which consists of two stages: Coarse-Grained Modal Pre-training and Fine-Grained Modal Fine-tuning. In the Pre-training phase, we introduce a Self-Supervised Similarity Contrast Learning method to facilitate effective knowledge transfer. During the Fine-tuning phase, we employ a Multi-Task Knowledge Distillation Fine-tuning Network, leveraging a teacher model to generate pseudo-labels and incorporating auxiliary tasks to enhance knowledge extraction from multimodal data, ultimately improving MNER performance. Experimental results demonstrate that MKD outperforms existing methods on the Twitter2015 and Twitter2017 datasets, achieving state-of-the-art results. Additionally, our findings indicate that knowledge can be effectively transferred between modalities, and simultaneous multitasking learning further boosts MNER performance.