N-Gram approach to prepare Crime-related Legal DataSet: A roadmap to classify Legal Text
摘要
While analyzing legal documents, legal professionals usually explore vital information in any legal case documents. Mostly, they manually extract major information from legal documents, which is time-consuming. Artificial Intelligence-based techniques can reduce this time-consuming nature. Dataset plays a major role in preparing any Artificial Intelligence based system in any application domain. Developing an Artificial Intelligence-based system for legal professionals requires knowledge of features that can be provided as a dataset. Hence in this paper, the authors have proposed an approach using NLP-based techniques like N-gram, Wordnet, Lemmatization, etc., on legal documents to prepare a legal dataset. As a case study, this research focused on dowry death cases, one of our society’s alarming crimes. The proposed methodology to prepare datasets has been determined to capture the major concepts of dowry death cases with more than 95% precision. In the future, a good number of legal datasets can be prepared on other woman-centric crimes using the proposed methodology.