Ordinal and Position Enhance the Framework of the Multimodal Dialogue System
摘要
Multimodal dialogue systems aim to process multimodal input information such as text and images simultaneously and then generate coherent and meaningful output dialogue responses. Although current methods have achieved notable progress, they have shortcomings in the following aspects: 1) Insufficient modeling of relationships between multimodal semantic elements, especially ignoring the interaction between the ordinal information in text and the position information of images. 2) Ineffectively integrating information from various modalities in multimodal conversations. To address these limitations, this paper proposes a framework for a multimodal dialogue system. We integrate both ordinal words from text and the position information of images as ordinal information to strengthen the interaction between multimodal semantic elements and to improve the understanding of user intent. In particular, in this framework, we incorporated the self-attention mechanism into the multimodal factorized bilinear pooling (MFB) method, which captures long-range dependencies between modalities, enhancing the effectiveness of multimodal data fusion. Finally, we conducted comprehensive experiments on public multimodal dialogue datasets (MMConv and MMD) and demonstrated that our proposed framework outperforms other methods in the experiments.