<p>Multimodal Named Entity Recognition (MNER) is essential for effective information extraction, yet traditional methods encounter several challenges. These include mismatches between text and images, insufficient utilization of image data, neglect of critical features, and difficulties in aligning semantic levels across modalities. Typically, these approaches focus on aligning and fusing multimodal data without integrating external knowledge. To address these shortcomings, we propose a <b>M</b>ulti-Granularity <b>K</b>nowledge <b>D</b>istillation model for MNER (MKD), which consists of two stages: Coarse-Grained Modal Pre-training and Fine-Grained Modal Fine-tuning. In the Pre-training phase, we introduce a Self-Supervised Similarity Contrast Learning method to facilitate effective knowledge transfer. During the Fine-tuning phase, we employ a Multi-Task Knowledge Distillation Fine-tuning Network, leveraging a teacher model to generate pseudo-labels and incorporating auxiliary tasks to enhance knowledge extraction from multimodal data, ultimately improving MNER performance. Experimental results demonstrate that MKD outperforms existing methods on the Twitter2015 and Twitter2017 datasets, achieving state-of-the-art results. Additionally, our findings indicate that knowledge can be effectively transferred between modalities, and simultaneous multitasking learning further boosts MNER performance.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Enhancing multimodal named entity recognition with multi-granularity knowledge distillation

  • Xinyu He,
  • Shixin Li,
  • Binhe Li

摘要

Multimodal Named Entity Recognition (MNER) is essential for effective information extraction, yet traditional methods encounter several challenges. These include mismatches between text and images, insufficient utilization of image data, neglect of critical features, and difficulties in aligning semantic levels across modalities. Typically, these approaches focus on aligning and fusing multimodal data without integrating external knowledge. To address these shortcomings, we propose a Multi-Granularity Knowledge Distillation model for MNER (MKD), which consists of two stages: Coarse-Grained Modal Pre-training and Fine-Grained Modal Fine-tuning. In the Pre-training phase, we introduce a Self-Supervised Similarity Contrast Learning method to facilitate effective knowledge transfer. During the Fine-tuning phase, we employ a Multi-Task Knowledge Distillation Fine-tuning Network, leveraging a teacher model to generate pseudo-labels and incorporating auxiliary tasks to enhance knowledge extraction from multimodal data, ultimately improving MNER performance. Experimental results demonstrate that MKD outperforms existing methods on the Twitter2015 and Twitter2017 datasets, achieving state-of-the-art results. Additionally, our findings indicate that knowledge can be effectively transferred between modalities, and simultaneous multitasking learning further boosts MNER performance.