Enhancing visual question answering with common sense knowledge: a data-driven neurosymbolic graph routing approach
摘要
Visual question answering (VQA) requires models to comprehend both visual and textual inputs, often necessitating multi-hop reasoning and external knowledge beyond image content. Despite recent advances, current VQA models struggle with complex questions that require reasoning over both structured knowledge and unstructured visual information. To address these limitations, we propose the neurosymbolic graph routing network (NeSyGRN), which integrates a graph routing network for stepwise reasoning with an enrichment mechanism leveraging the common sense knowledge graph (CSKG). This approach enriches scene graphs and question representations with common sense knowledge, enhancing the reasoning capabilities of VQA through structured data augmentation. NeSyGRN is evaluated on benchmark datasets, demonstrating substantial improvements over state-of-the-art baselines. On the KR-VQA dataset, NeSyGRN achieves an overall accuracy of 74.52%, surpassing prior best results by 8%. Ablation studies validate the critical role of knowledge enrichment, with scene graph enrichment contributing most significantly to the improvements. On the GQA dataset, NeSyGRN achieves an accuracy of 83.43%, setting new state-of-the-art results in consistency, plausibility, and contextual validity. These promising results highlight the effectiveness of integrating structured common sense knowledge for VQA, reinforcing its potential for applications in data-rich environments such as health care, education, and decision support systems.