Multimodal Large Language Models (MLLMs) have recently been very successful in various tasks including handwriting recognition (HWR), Visual Question-Answering (VQA), object detection and classification. However, there has not been any effort to evaluate these models with images from the legal domain. In this work, we present the HIFIRE (Handwritten Indian First Information Reports in English) dataset containing FIR document images from different police stations in India. These FIR images are diverse and contain both printed field-names and hand-written texts, making this a challenging dataset. The dataset is divided into two parts – (1) HIFIRE-HWR containing 20,078 manually annotated images for handwriting recognition, and (2) HIFIRE-Doc containing 543 annotated document images for three tasks – (i) Text Object Detection and Classification, (ii) Document Visual Question Answering, and (iii) Legal Reasoning with Visual Question Answering. We apply various MLLMs on the HIFIRE datasets to benchmark their performance on the aforementioned tasks. The moderate performances of most MLLMs show that the legal domain-specific tasks are challenging and need better models. We also show that our dataset is suitable for fine-tuning MLLMs to help achieve substantial improvements on the tasks.