Knowledge-based visual question answering (VQA) requires external knowledge in addition to the image content to answer questions. Recent studies convert images to text descriptions and then generate answers or acquire implicit knowledge using a large language model (LLM). These methods achieve encouraging results with the strong knowledge retrieval and reasoning capabilities of LLMs. However, methods that incorporate LLMs are limited by the discrepancies between images and their text descriptions presented to LLMs. To address this challenge, we present RAVL, a retrieval-augmented visual language model (VLM) framework for knowledge-based VQA. Specifically, we first fine-tune a VLM on the knowledge-based VQA task with inputs consisting of retrieved knowledge and image-question pairs to adapt the VLM to inputs with retrieved knowledge. After that, we adapt the retrieval module to the fine-tuned VLM using supervision signals provided by the VLM, enabling the retrieved knowledge to improve the VLM perplexity. RAVL overcomes the limitation of visual information loss and improves the effectiveness of VLMs with external knowledge. We conduct experiments on OK-VQA dataset and our method achieves 65.73% accuracy, surpassing the previous state-of-the-art method (+3.63%).

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

RAVL: A Retrieval-Augmented Visual Language Model Framework for Knowledge-Based Visual Question Answering

  • Naiquan Chai,
  • Dongsheng Zou,
  • Jiyuan Liu,
  • Hao Wang,
  • Yuming Yang,
  • Xinyi Song

摘要

Knowledge-based visual question answering (VQA) requires external knowledge in addition to the image content to answer questions. Recent studies convert images to text descriptions and then generate answers or acquire implicit knowledge using a large language model (LLM). These methods achieve encouraging results with the strong knowledge retrieval and reasoning capabilities of LLMs. However, methods that incorporate LLMs are limited by the discrepancies between images and their text descriptions presented to LLMs. To address this challenge, we present RAVL, a retrieval-augmented visual language model (VLM) framework for knowledge-based VQA. Specifically, we first fine-tune a VLM on the knowledge-based VQA task with inputs consisting of retrieved knowledge and image-question pairs to adapt the VLM to inputs with retrieved knowledge. After that, we adapt the retrieval module to the fine-tuned VLM using supervision signals provided by the VLM, enabling the retrieved knowledge to improve the VLM perplexity. RAVL overcomes the limitation of visual information loss and improves the effectiveness of VLMs with external knowledge. We conduct experiments on OK-VQA dataset and our method achieves 65.73% accuracy, surpassing the previous state-of-the-art method (+3.63%).