<p>Knowledge-based visual question answering has made progress in cross-modal feature fusion and external knowledge integration, yet it still suffers from limitations in dynamic knowledge updating and continual learning, which constrain knowledge coverage and noise suppression. To address these issues, we propose a visual question answering model based on autonomous multi-modal knowledge learning, in which knowledge learning is divided into two stages: knowledge accumulation and knowledge updating. Specifically, large-scale multi-modal knowledge is first accumulated by constructing pseudo data, and then dynamically filtered and iteratively updated through an autonomous updating strategy to reduce noise and improve knowledge quality. During inference, the model actively retrieves and invokes relevant knowledge via a vector index, enabling automatic knowledge utilization and continual accumulation. Compared with the baseline models MuKEA and CMLR, the proposed method achieves accuracy improvements of 4.21% and 2.87% on the OK-VQA dataset, respectively, demonstrating that it effectively alleviates insufficient knowledge coverage and noise sensitivity in knowledge-based VQA and significantly enhances cross-modal understanding.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Visual question answering model based on multi-modal knowledge autonomous learning

  • Yuanlong Wang,
  • Zhiru Xu,
  • Houshuai Wang,
  • Zhiwei Wu,
  • Hu Zhang

摘要

Knowledge-based visual question answering has made progress in cross-modal feature fusion and external knowledge integration, yet it still suffers from limitations in dynamic knowledge updating and continual learning, which constrain knowledge coverage and noise suppression. To address these issues, we propose a visual question answering model based on autonomous multi-modal knowledge learning, in which knowledge learning is divided into two stages: knowledge accumulation and knowledge updating. Specifically, large-scale multi-modal knowledge is first accumulated by constructing pseudo data, and then dynamically filtered and iteratively updated through an autonomous updating strategy to reduce noise and improve knowledge quality. During inference, the model actively retrieves and invokes relevant knowledge via a vector index, enabling automatic knowledge utilization and continual accumulation. Compared with the baseline models MuKEA and CMLR, the proposed method achieves accuracy improvements of 4.21% and 2.87% on the OK-VQA dataset, respectively, demonstrating that it effectively alleviates insufficient knowledge coverage and noise sensitivity in knowledge-based VQA and significantly enhances cross-modal understanding.