Visual Question Answering (VQA) is a challenging task in natural language processing and computer vision that requires sophisticated reasoning over the visual content to provide an accurate answer to a question. This paper proposes an efficient deep learning framework for Malayalam Visual Question Answering (MVQA) that can answer a specific natural language question in Malayalam about an image. A Malayalam Visual Question Answering (MVQA) dataset was created by translating English question answer pairs from the Visual Genome dataset. By extracting the object-level spatial information from contextual parts of the image using FR-CNN (Faster R-CNN), the proposed model predicts a vernacular answer from multiple answers. The proposed MVQA framework comprises (1) a region-based convolutional neural network (CNN) for visual feature extraction, (2) Long Short-Term Memory cells (LSTMs) for language modeling, and (3) multimodal fusion for integrating visual and language representations. We evaluate the proposal for visual question answering using the proposed framework, and experimental results on the self-created Malayalam VQA dataset show the effectiveness of this framework for Malayalam VQA.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Object Aware Visual Question Answering Model for Malayalam

  • Abhishek Gopinath Kovath,
  • O. K. Sikha

摘要

Visual Question Answering (VQA) is a challenging task in natural language processing and computer vision that requires sophisticated reasoning over the visual content to provide an accurate answer to a question. This paper proposes an efficient deep learning framework for Malayalam Visual Question Answering (MVQA) that can answer a specific natural language question in Malayalam about an image. A Malayalam Visual Question Answering (MVQA) dataset was created by translating English question answer pairs from the Visual Genome dataset. By extracting the object-level spatial information from contextual parts of the image using FR-CNN (Faster R-CNN), the proposed model predicts a vernacular answer from multiple answers. The proposed MVQA framework comprises (1) a region-based convolutional neural network (CNN) for visual feature extraction, (2) Long Short-Term Memory cells (LSTMs) for language modeling, and (3) multimodal fusion for integrating visual and language representations. We evaluate the proposal for visual question answering using the proposed framework, and experimental results on the self-created Malayalam VQA dataset show the effectiveness of this framework for Malayalam VQA.