A Novel Preprocessing Method for Transforming Federal Sentencing Data to Ensure Unbiased AI Adjudication Research Using Large Language Models
摘要
In an effort to create objective AI for judicial decision-making, this study offers a unique preprocessing method for the JUSTFAIR data-set of federal criminal sentencing records, which will be used in Large Language Models (LLMs). While LLMs struggle with numerical and categorical data, traditional machine learning methods struggle with the complex, linked nature of judicial data. Our approach tackles these problems in a multi-phase manner: (1) data cleansing with MongoDB; (2) textual categorization of numerical data using USSC and JUSTFAIR codebooks; (3) enlarging legal acronyms; and (4) narrative story creation from discrete data points with Claude 3.5 Sonnet. LLMs can now more fully understand the complex relationships involved in sentencing decisions because to this shift. Our method establishes the groundwork for LLMs to be trained on less biased court data, which could transform AI applications in the legal space. The work offers ramifications outside of the legal system to other industries with sensitive, complicated data, contributing to AI ethics and bias reduction. In this paper, we present a detailed exposition of our methodological framework and offer a comprehensive evaluation of outcomes, thereby establishing the credibility and significance of our findings.