Mgsc: multimodal generation and self-supervised contrast learning for mitigating language bias in visual question answering
摘要
Visual Question Answering (VQA) models often suffer from language bias issues, which refer to the tendency of models to generate answers solely based on superficial correlations in the question-answer pairs from the training set, without properly understanding the visual information. To mitigate language bias, prior studies have primarily employed auxiliary question-only models or attempted dataset rebalancing. However, these methods frequently overlook the full potential of visual and linguistic information, thus neglecting the importance of capturing complementary relationships between the two modalities. To address this issue, we propose Multimodal Generation and Self-Supervised Contrast Learning (MGSC), which first leverages a generative adversarial network to train a bias model, guiding the target model to capture complementary information between vision and language. The proposed method then generates counterfactual samples by shuffling question-image pairs, encouraging the model to focus on visual semantics and reducing its reliance on superficial correlations between question and answer patterns. The effectiveness of our method is validated through experiments conducted on the VQA-CP v2, VQA v2, and VQA-CE datasets, highlighting its robust generalization across diverse evaluation scenarios.