Logic-based interpretability and analysis of neural legal AI models
摘要
The growing adoption of artificial intelligence in judicial applications exposes a critical limitation of neural network-based legal AI systems: their decision-making processes are inherently opaque, which undermines the reliability of their outcomes, especially in cases involving personal liberty. Existing interpretability methods primarily rely on instance-based analysis and local input-output attributions; however, they are unable to capture the models’ global decision-making mechanisms, which greatly limits their applicability in legal contexts where both interpretability and predictive accuracy are essential. To address this challenge, this work proposes a logic-based interpretability and analysis framework that integrates the neural network prediction process with formal logical representations to construct transparent decision traces, capturing the relationships between legal elements and model outputs. Based on this representation, we introduce a method for evaluating the logical relationship between the extracted model logic and statutory legal elements, define the concept of conditional satisfiability in the legal domain, and apply influence-based metrics to quantify the importance of legal elements across different cases. Our analysis of the influence rankings reveals significant inconsistencies between the internal decision logic of legal AI systems and actual statutory logic, confirming that even highly accurate models may fail to align with legal decision-making principles. Our framework offers a globally interpretable analysis methodology for inspecting neural legal AI models, thereby supporting risk identification and trustworthiness assessment in legal AI applications, providing a foundation for trustworthy and interpretable AI in law.