This paper presents a novel approach to enhancing Visual Question Answering (VQA) systems in the Vietnamese language by integrating La rge Language Models (LLMs) with Optical Character Recognition (OCR) technology. VQA systems are designed to understand and answer questions based on visual content, but their performance may not be good when dealing with textual information embedded in images. To address this challenge, we propose a hybrid framework that leverages the strengths of LLMs in natural language understanding and generation, alongside OCR systems for accurate text extraction from images. Our approach involves the preprocessing of images through OCR to extract textual data, which is then combined with visual features and processed by an LLM to generate more accurate and contextually relevant answers. We conducted extensive experiments on a Vietnamese VQA dataset to evaluate the effectiveness of our method. The results demonstrate significant improvements in the accuracy of the answer and contextual understanding, with a 2% increase in the BLEU score and a 1.11 increase in the CIDEr score compared to the state-of-the-art VQA models.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Enhancing Visual Question Answering in Vietnamese Using Large Language Models Combined with OCR Systems

  • Nguyen Van Nha,
  • Phung The Huan,
  • Le Minh Tuan,
  • Le Hoang Son

摘要

This paper presents a novel approach to enhancing Visual Question Answering (VQA) systems in the Vietnamese language by integrating La rge Language Models (LLMs) with Optical Character Recognition (OCR) technology. VQA systems are designed to understand and answer questions based on visual content, but their performance may not be good when dealing with textual information embedded in images. To address this challenge, we propose a hybrid framework that leverages the strengths of LLMs in natural language understanding and generation, alongside OCR systems for accurate text extraction from images. Our approach involves the preprocessing of images through OCR to extract textual data, which is then combined with visual features and processed by an LLM to generate more accurate and contextually relevant answers. We conducted extensive experiments on a Vietnamese VQA dataset to evaluate the effectiveness of our method. The results demonstrate significant improvements in the accuracy of the answer and contextual understanding, with a 2% increase in the BLEU score and a 1.11 increase in the CIDEr score compared to the state-of-the-art VQA models.