RDA: Regularized Domain Adaptation for Multimedia Event Extraction
摘要
Multimedia event extraction focuses on identifying structured events and arguments from multimedia documents. Due to the scarcity of parallel textual-visual events, most methods rely on unlabeled image-caption pairs or synthetic data, often suffering from label shifts caused by domain discrepancy. To address this, we propose RDA, a Regularized Domain Adaptation framework for multimedia event extraction, which uses a coarse-to-fine domain adaptation method. In the coarse-grained phase, we use a dual-encoder architecture and multimodal fusion module to learn unified cross-modal representations. In the fine-grained phase, we apply regularization to align domain discrepancy based on global image features, improving model generalization. Experiments on the M2E2 benchmark show that RDA achieves state-of-the-art performance, significantly improving both visual and multimedia event extraction.