Knowledge-based Visual Question Answering (KB-VQA) is a crucial issue act evaluation metrics undermine model performance. To address these issut the intersection of computer vision and natural language processing. Despite the recent emergence of large language models (LLMs) that have spurred rapid advancements in the KB-VQA field, several challenges persist: (1) Inadequate visual representation leads to the loss of critical visual information, resulting in unreliable reasoning; (2) Low-quality external knowledge interferes with answer predictions; (3) Stries, we first enrich visual representations by generating plentiful image captions, then integrate visual and textual information to obtain high-quality, multi-source external knowledge, and finally employ simple alignment strategies to enhance the robustness of answer formats. Importantly, all components within our methodological framework are interchangeable, rendering our framework more flexible and resource-efficient. Extensive experiments on the dataset demonstrate that our approach achieves state-of-the-art performance.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Prompting Large Models for Knowledge and Reasoning Augmentation in KB-VQA

  • Qiang Liu,
  • Mengxi Ying,
  • Peng Xiao,
  • Gan Li,
  • Xinpan Yuan

摘要

Knowledge-based Visual Question Answering (KB-VQA) is a crucial issue act evaluation metrics undermine model performance. To address these issut the intersection of computer vision and natural language processing. Despite the recent emergence of large language models (LLMs) that have spurred rapid advancements in the KB-VQA field, several challenges persist: (1) Inadequate visual representation leads to the loss of critical visual information, resulting in unreliable reasoning; (2) Low-quality external knowledge interferes with answer predictions; (3) Stries, we first enrich visual representations by generating plentiful image captions, then integrate visual and textual information to obtain high-quality, multi-source external knowledge, and finally employ simple alignment strategies to enhance the robustness of answer formats. Importantly, all components within our methodological framework are interchangeable, rendering our framework more flexible and resource-efficient. Extensive experiments on the dataset demonstrate that our approach achieves state-of-the-art performance.