<p>Identifying salient sentences in long judicial decisions is a major challenge in legal document understanding. This process is time-consuming and difficult even for legal experts. In countries like India, where the common law system is heavily based on past cases, quickly recognizing key sentences is important for rapid resolution of cases. The emergence of deep learning models for language processing tasks can help legal professionals make faster and more informed decisions. However, using these models to automate saliency detection requires annotated data for training and evaluation, which is costly and requires trained legal professionals. Here we introduce the task of saliency detection in Indian judicial decisions, along with the curation of a concise gold standard dataset tailored to this task. An adaptive data augmentation strategy is introduced to address the scarcity of labeled data. This strategy dynamically adjusts the amount of augmentation applied to training samples based on the model’s current classification performance on the validation set. By monitoring performance at different levels, the augmentation process can be fine-tuned, allowing model training to be stopped with the optimal augmented dataset. Legal experts provide a list of protected terms to ensure that specific legal terms remain unchanged during the augmentation process. The methodology focuses on training a deep learning-based Convolutional Bidirectional LSTM model, evaluating the performance with and without augmentation, and conducting a comparative analysis against fine-tuned transformer models used as baselines. The results demonstrate improvements in model performance while preserving protected terms and carefully managing augmentation levels. Additionally, the trained deep learning models on the augmented set proved to be more resource-efficient compared to fine-tuned models.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Adaptive data augmentation for salient sentence identification in Indian judicial decisions

  • Reshma Sheik,
  • Krishnadas Nair,
  • S. K. Manu Krishna,
  • S. Jaya Nirmala

摘要

Identifying salient sentences in long judicial decisions is a major challenge in legal document understanding. This process is time-consuming and difficult even for legal experts. In countries like India, where the common law system is heavily based on past cases, quickly recognizing key sentences is important for rapid resolution of cases. The emergence of deep learning models for language processing tasks can help legal professionals make faster and more informed decisions. However, using these models to automate saliency detection requires annotated data for training and evaluation, which is costly and requires trained legal professionals. Here we introduce the task of saliency detection in Indian judicial decisions, along with the curation of a concise gold standard dataset tailored to this task. An adaptive data augmentation strategy is introduced to address the scarcity of labeled data. This strategy dynamically adjusts the amount of augmentation applied to training samples based on the model’s current classification performance on the validation set. By monitoring performance at different levels, the augmentation process can be fine-tuned, allowing model training to be stopped with the optimal augmented dataset. Legal experts provide a list of protected terms to ensure that specific legal terms remain unchanged during the augmentation process. The methodology focuses on training a deep learning-based Convolutional Bidirectional LSTM model, evaluating the performance with and without augmentation, and conducting a comparative analysis against fine-tuned transformer models used as baselines. The results demonstrate improvements in model performance while preserving protected terms and carefully managing augmentation levels. Additionally, the trained deep learning models on the augmented set proved to be more resource-efficient compared to fine-tuned models.