Multimedia event extraction focuses on identifying structured events and arguments from multimedia documents. Due to the scarcity of parallel textual-visual events, most methods rely on unlabeled image-caption pairs or synthetic data, often suffering from label shifts caused by domain discrepancy. To address this, we propose RDA, a Regularized Domain Adaptation framework for multimedia event extraction, which uses a coarse-to-fine domain adaptation method. In the coarse-grained phase, we use a dual-encoder architecture and multimodal fusion module to learn unified cross-modal representations. In the fine-grained phase, we apply regularization to align domain discrepancy based on global image features, improving model generalization. Experiments on the M2E2 benchmark show that RDA achieves state-of-the-art performance, significantly improving both visual and multimedia event extraction.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

RDA: Regularized Domain Adaptation for Multimedia Event Extraction

  • Yuhui Zhang,
  • Yongxiu Xu,
  • Minghao Tang,
  • Xinkui Lin,
  • Yubin Wang,
  • Hongbo Xu,
  • Gaopeng Gou

摘要

Multimedia event extraction focuses on identifying structured events and arguments from multimedia documents. Due to the scarcity of parallel textual-visual events, most methods rely on unlabeled image-caption pairs or synthetic data, often suffering from label shifts caused by domain discrepancy. To address this, we propose RDA, a Regularized Domain Adaptation framework for multimedia event extraction, which uses a coarse-to-fine domain adaptation method. In the coarse-grained phase, we use a dual-encoder architecture and multimodal fusion module to learn unified cross-modal representations. In the fine-grained phase, we apply regularization to align domain discrepancy based on global image features, improving model generalization. Experiments on the M2E2 benchmark show that RDA achieves state-of-the-art performance, significantly improving both visual and multimedia event extraction.