<p>The proliferation of multilingual Visual Question Answering (VQA) datasets is paramount for augmenting the capabilities of large language models (LLMs) and multi-modal LLMs, thereby enabling them to adeptly capture the intricate linguistic subtleties and visual complexities inherent across diverse languages. This scholarly article delineates <b>HW-MLVQA</b>, a novel dataset meticulously crafted to mitigate the dearth of genuine handwritten datasets pivotal for multilingual document comprehension. <b>HW-MLVQA</b> encompasses an extensive collection of 14,000 handwritten images complemented by 24,000 question-answer pairs. Furthermore, HW-MLVQA provides a robust benchmark evaluation framework spanning three distinct modalities: text, image, and an integrated image &amp; text modality. The dataset also facilitates a rigorous assessment of proprietary and open-source Optical Character Recognition (OCR) systems to simulate a real-world scenario where ground-truth transcriptions are inaccessible. <b>HW-MLVQA</b> aspires to catalyze transformative advancements in multilingual handwritten document interpretation, fostering innovation and scholarly inquiry within this specialized domain. The dataset and code-base will be made publicly available upon acceptance.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

HW-MLVQA: a novel handwritten multilingual dataset for visual question answering and evaluation

  • Aniket Pal,
  • Ajoy Mondal,
  • C. V. Jawahar

摘要

The proliferation of multilingual Visual Question Answering (VQA) datasets is paramount for augmenting the capabilities of large language models (LLMs) and multi-modal LLMs, thereby enabling them to adeptly capture the intricate linguistic subtleties and visual complexities inherent across diverse languages. This scholarly article delineates HW-MLVQA, a novel dataset meticulously crafted to mitigate the dearth of genuine handwritten datasets pivotal for multilingual document comprehension. HW-MLVQA encompasses an extensive collection of 14,000 handwritten images complemented by 24,000 question-answer pairs. Furthermore, HW-MLVQA provides a robust benchmark evaluation framework spanning three distinct modalities: text, image, and an integrated image & text modality. The dataset also facilitates a rigorous assessment of proprietary and open-source Optical Character Recognition (OCR) systems to simulate a real-world scenario where ground-truth transcriptions are inaccessible. HW-MLVQA aspires to catalyze transformative advancements in multilingual handwritten document interpretation, fostering innovation and scholarly inquiry within this specialized domain. The dataset and code-base will be made publicly available upon acceptance.