Automatic Detection of Errors in LLM Large Benchmarks Using Frontier Model Consensus
摘要
The rapid advancement of Large Language Models (LLMs) has led to their widespread adoption in various academic and business applications. However, the reliability of these models remains a concern, particularly in situations where their outputs cannot be fully trusted. This paper presents an approach to identify potential errors in large LLM benchmarks by leveraging the consensus of frontier models. Our study focuses on the Massive Multitask Language Understanding (MMLU) benchmark, a popular dataset used to evaluate the performance of LLMs across a wide range of subjects. Our approach demonstrates the potential for using model consensus as a tool to detect benchmark errors and can lead to the creation of cleaner, more accurate datasets.