There is a need for a VQA framework which is semantics oriented and compliant with the semantic standards of the Web 3.0, this paper proposes a VQA framework that encompasses semantics-oriented learning – reasoning, where the datasets of questions and images are separately maintained wherein the informative terms and keywords are extracted from both the datasets and XGBoost classifier are subjected to classification of the dataset of questions and dataset of images. Semantics oriented language attenuation is achieved through auxiliary knowledge inclusion using crawled web data, wiki data API, name effigy recognition is also encompassed which is further subjected through generation of metadata in order to increase the overall density of the auxiliary knowledge and reduce the gap between the knowledge of external web and the knowledge that moves into the model. The classification of the metadata is achieved using autoencoders classifiers, semantics-oriented similarity and reasoning through semantic similarity is achieved using normalized compression distance and lance William index with a differential threshold and step deviance methods to yield the best-in-class associations of questions and images for visual question answering. A precision of 95.45% and F-Measure of 95.25% has been robust with FDR of 0.05 has been achieved through the proposed model.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

SOVQA: Semantics Oriented Framework for Visual Question Answering Aggregating Regulated Knowledge and Partial Learning

  • Abhinav Hebbar,
  • Gerard Deepak

摘要

There is a need for a VQA framework which is semantics oriented and compliant with the semantic standards of the Web 3.0, this paper proposes a VQA framework that encompasses semantics-oriented learning – reasoning, where the datasets of questions and images are separately maintained wherein the informative terms and keywords are extracted from both the datasets and XGBoost classifier are subjected to classification of the dataset of questions and dataset of images. Semantics oriented language attenuation is achieved through auxiliary knowledge inclusion using crawled web data, wiki data API, name effigy recognition is also encompassed which is further subjected through generation of metadata in order to increase the overall density of the auxiliary knowledge and reduce the gap between the knowledge of external web and the knowledge that moves into the model. The classification of the metadata is achieved using autoencoders classifiers, semantics-oriented similarity and reasoning through semantic similarity is achieved using normalized compression distance and lance William index with a differential threshold and step deviance methods to yield the best-in-class associations of questions and images for visual question answering. A precision of 95.45% and F-Measure of 95.25% has been robust with FDR of 0.05 has been achieved through the proposed model.