LegSegSC: A Silver-Standard Rhetorical Role Labeled Dataset of Indian Supreme Court Judgments
摘要
Legal documents, such as court judgments and legal briefs, are complex, lengthy, and challenging to analyze manually. The increasing digitization of legal resources has created a growing need for automated NLP-based systems to streamline document review, rhetorical role (RR) labeling, and summarization. However, manual annotation of RR-labels is time-consuming, expensive, and requires domain expertise, making large-scale labeled datasets difficult to obtain. To reduce manual annotation efforts, we propose a model-assisted approach to generate a silver-standard dataset using a gold-standard dataset. We fine-tuned multiple pre-trained models, including DistillBERT, LegalBERT, Legal RoBERTa, ERNIE 2.0, and SCI-BERT, on the human-annotated BUILDNyAI dataset. Legal RoBERTa achieved the highest accuracy of 71.7%, which we used to create LegSegSC, a dataset of 5472 Indian Supreme Court judgments with 1.1 million automatically labeled sentences. Our contributions include: (1) fine-tuning multiple models for RR classification in legal texts, (2) achieving state-of-the-art performance with Legal RoBERTa, (3) introducing a silver-standard large-scale dataset for future research, and (4) demonstrating the feasibility of reducing manual annotation effort while maintaining high-quality labels.