Discovering Inconsistencies in Documents with Long-Context LLMs
摘要
The increasing complexity and scale of technical document corpora present challenges for consistency verification, particularly in politically sensitive or high-stakes contexts. This paper proposes an iterative approach that integrates long-context large language models (LLMs), human expertise, and hybrid clustering mechanisms to address these challenges. The approach focuses on two types of inconsistencies: real inconsistencies, such as contradictory statements or omissions, and fabricated inconsistencies, which are plausible yet artificially introduced. This paper uses the Swiss National Cooperative for the Disposal of Radioactive Waste (Nagra) and its corpus of up to 300 technical documents as a case study. Experimental results suggest that targeted structuring of document contexts improves recall in inconsistency detection. The findings highlight the potential of combining structured human input with LLM-based reasoning for improving document integrity and trustworthiness. Future work will focus on refining the approach, including automated clustering strategies and optimization of prompt engineering.