<p>Indian court judgment reports frequently include complicated words and sentences, making it difficult for the general public and legal experts to understand these documents. Legal organizations hire legal experts to summarize complex and lengthy legal texts. Hence, a variety of techniques have been created to construct the summaries. This study investigates the application of InLegalBERT, a pre-trained legal language model, for summarizing Indian legal documents. We propose a novel framework that utilizes InLegalBERT’s sentence embeddings combined with K-Means clustering to extract and prioritize legally significant information for generating concise summaries. Unlike traditional methods, our approach integrates domain-specific knowledge to enhance the accuracy and relevance of summaries. The framework was evaluated at compression ratios of 10%, 20%, and 30%, and benchmarked against five models: Legal Pegasus, T5 base, BART, BERT, and ChatGPT. At a 30% compression ratio, our method outperformed others with a ROUGE-L F1 score of 0.3858, precision of 0.3585, recall of 0.4526, and a beta score of 0.4481. Additionally, qualitative evaluation using Krippendorff’s Alpha confirmed the consistency and reliability of the summaries.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Advancements in legal text summarization: integrating InLegalBERT for effective extractive summarization

  • Saloni Sharma,
  • Piyush Pratap Singh

摘要

Indian court judgment reports frequently include complicated words and sentences, making it difficult for the general public and legal experts to understand these documents. Legal organizations hire legal experts to summarize complex and lengthy legal texts. Hence, a variety of techniques have been created to construct the summaries. This study investigates the application of InLegalBERT, a pre-trained legal language model, for summarizing Indian legal documents. We propose a novel framework that utilizes InLegalBERT’s sentence embeddings combined with K-Means clustering to extract and prioritize legally significant information for generating concise summaries. Unlike traditional methods, our approach integrates domain-specific knowledge to enhance the accuracy and relevance of summaries. The framework was evaluated at compression ratios of 10%, 20%, and 30%, and benchmarked against five models: Legal Pegasus, T5 base, BART, BERT, and ChatGPT. At a 30% compression ratio, our method outperformed others with a ROUGE-L F1 score of 0.3858, precision of 0.3585, recall of 0.4526, and a beta score of 0.4481. Additionally, qualitative evaluation using Krippendorff’s Alpha confirmed the consistency and reliability of the summaries.