Reading Indian legal texts is often exhaustive. Indian case documents are usually less organized and have more errors than those from other countries. This study aims to help people quickly understand large legal documents. We created a new dataset with 10,000 judgments from the Supreme Court of India, along with their handwritten summaries. The dataset is cleaned to fix legal abbreviations, punctuation errors, and ensure proper sentence structure. Each judgment is annotated with attributes such as case ID, date of judgment, names of the plaintiff and defendants, judge who delivered the final verdict, cited acts, citations, main judgment, and its corresponding headnote. In the results section, we provide statistical analyses of the judgments and their headnotes, offering useful insights for future research. Beyond legal document summarization, potential applications of this dataset include information retrieval, citation analysis, and predicting decisions made by specific judges.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Structuring and Text Summarization of Indian Legal Documents

  • Pawan Kumar,
  • Bablu Kumar,
  • Pradeepika Verma,
  • Anshul Verma

摘要

Reading Indian legal texts is often exhaustive. Indian case documents are usually less organized and have more errors than those from other countries. This study aims to help people quickly understand large legal documents. We created a new dataset with 10,000 judgments from the Supreme Court of India, along with their handwritten summaries. The dataset is cleaned to fix legal abbreviations, punctuation errors, and ensure proper sentence structure. Each judgment is annotated with attributes such as case ID, date of judgment, names of the plaintiff and defendants, judge who delivered the final verdict, cited acts, citations, main judgment, and its corresponding headnote. In the results section, we provide statistical analyses of the judgments and their headnotes, offering useful insights for future research. Beyond legal document summarization, potential applications of this dataset include information retrieval, citation analysis, and predicting decisions made by specific judges.