The goal of the Indian legal system is to provide fair access to justice, but the use of complex legal language makes it difficult for the public and small businesses to understand. This study aims to address this issue by introducing a novel approach using the Large Language Model (LLM) to simplify legal terminology in India. We use a priorly trained Bidirectional Encoder Representation Transformers (BERT), specifically the BERT base uncased version, to bridge the gap between legal jargon and mutual understanding. Enhance the model’s effectiveness in the legal domain, we pre-train it on a large legal corpus of data of fourteen million legal terms extracted from Indian legal documents. Additionally, we conduct further training on a dataset holding 59,770 legal terms and their simpler alternatives sourced from the official Indian Law and Justice website. Various data pre-processing techniques such as tokenization, padding, and augmentation methods including synonym substitution, back-translation, and random reordering are employed to improve the model’s robustness and applicability. The use of mixed precision training with gradient scaling perfects the training process within the Transformer model architecture. Through simplifying legal texts, this LLM approach enables individuals to better understand their legal rights, promoting greater access to justice within the Indian legal system. The Optical Character Recognition (OCR) and Fuzzy-Wuzzy string matching are used in this study, as they hold the potential for scanning and predicting relevant words in the given document, thereby processing, and enhancing its capabilities in addressing the variations in legal terminology.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Transformer-Powered Legal Jargon Simplification: An OCR-Based LLM Approach for Indian Law and Justice

  • J. Sachin Mathew,
  • R. Sivakami

摘要

The goal of the Indian legal system is to provide fair access to justice, but the use of complex legal language makes it difficult for the public and small businesses to understand. This study aims to address this issue by introducing a novel approach using the Large Language Model (LLM) to simplify legal terminology in India. We use a priorly trained Bidirectional Encoder Representation Transformers (BERT), specifically the BERT base uncased version, to bridge the gap between legal jargon and mutual understanding. Enhance the model’s effectiveness in the legal domain, we pre-train it on a large legal corpus of data of fourteen million legal terms extracted from Indian legal documents. Additionally, we conduct further training on a dataset holding 59,770 legal terms and their simpler alternatives sourced from the official Indian Law and Justice website. Various data pre-processing techniques such as tokenization, padding, and augmentation methods including synonym substitution, back-translation, and random reordering are employed to improve the model’s robustness and applicability. The use of mixed precision training with gradient scaling perfects the training process within the Transformer model architecture. Through simplifying legal texts, this LLM approach enables individuals to better understand their legal rights, promoting greater access to justice within the Indian legal system. The Optical Character Recognition (OCR) and Fuzzy-Wuzzy string matching are used in this study, as they hold the potential for scanning and predicting relevant words in the given document, thereby processing, and enhancing its capabilities in addressing the variations in legal terminology.