Understanding the Roles of Visual Modality in Multimodal Dialogue: An Empirical Study
摘要
Multimodal dialogue systems aim to simulate human-world interactions by perceiving information from various modalities. By developing effective strategies to integrate textual and visual modalities, multimodal dialogue models have yielded significant advancements. However, limited research has systematically and comprehensively explored the impact of the visual modality on dialogue response generation. To address this, this paper presents empirical studies conducted on a prominent multimodal dialogue model named Maria. To facilitate our analysis, we design a diagnostic approach utilizing cross-modal input ablation experiments. Our methodology involves manipulating the visual content of images by substituting visual regions or removing textual visual concepts while keeping the remaining inputs unchanged. Through experiments conducted on three datasets, we observe that the enhancements achieved through visual region features exhibit varying levels of stability across datasets. However, carefully introducing textual visual concepts, particularly those providing supplementary information, yields positive effects. Furthermore, discrepancies arise when constructing multimodal dialogue datasets using different approaches, underscoring the attention needed for future research in this field.