<p>Multimodal Large Language Models (MLLMs) have recently been very successful in various tasks including handwriting recognition (HWR), Visual Question-Answering (VQA), object detection and classification. However, there has not been any effort to evaluate these models with images from the legal domain. In this work, we present the <Emphasis FontCategory="NonProportional">HIFIRE</Emphasis> (Handwritten Indian First Information Reports in English) dataset containing FIR document images from different police stations in India. These FIR images are diverse and contain both printed field-names and hand-written texts, making this a challenging dataset. The dataset is divided into two parts – (1)&#xa0;<Emphasis FontCategory="NonProportional">HIFIRE</Emphasis>-HWR containing 20,078 manually annotated images for handwriting recognition, and (2)&#xa0;<Emphasis FontCategory="NonProportional">HIFIRE</Emphasis>-Doc containing 543 annotated document images for three tasks – (i)&#xa0;Text Object Detection and Classification, (ii)&#xa0;Document Visual Question Answering, and (iii)&#xa0;Legal Reasoning with Visual Question Answering. We apply various MLLMs on the <Emphasis FontCategory="NonProportional">HIFIRE</Emphasis> datasets to benchmark their performance on the aforementioned tasks. The moderate performances of most MLLMs show that the legal domain-specific tasks are challenging and need better models. We also show that our dataset is suitable for fine-tuning MLLMs to help achieve substantial improvements on the tasks.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

How well do MLLMs understand handwritten legal documents? A novel dataset for benchmarking

  • Sagar Chakraborty,
  • Gaurav Harit,
  • Saptarshi Ghosh

摘要

Multimodal Large Language Models (MLLMs) have recently been very successful in various tasks including handwriting recognition (HWR), Visual Question-Answering (VQA), object detection and classification. However, there has not been any effort to evaluate these models with images from the legal domain. In this work, we present the HIFIRE (Handwritten Indian First Information Reports in English) dataset containing FIR document images from different police stations in India. These FIR images are diverse and contain both printed field-names and hand-written texts, making this a challenging dataset. The dataset is divided into two parts – (1) HIFIRE-HWR containing 20,078 manually annotated images for handwriting recognition, and (2) HIFIRE-Doc containing 543 annotated document images for three tasks – (i) Text Object Detection and Classification, (ii) Document Visual Question Answering, and (iii) Legal Reasoning with Visual Question Answering. We apply various MLLMs on the HIFIRE datasets to benchmark their performance on the aforementioned tasks. The moderate performances of most MLLMs show that the legal domain-specific tasks are challenging and need better models. We also show that our dataset is suitable for fine-tuning MLLMs to help achieve substantial improvements on the tasks.