HalluScope: A Comprehensive Dataset for Evaluating Hallucination in Large Language Models Across Multiple Domains
摘要
This study presents HalluScope, a benchmark specifically developed for evaluating hallucinations in large language models. HalluScope comprises 800 adversarially designed questions spanning multiple domains, systematically categorized into selective, temporal, imitative, factual, and overconfidence hallucinations. The dataset was constructed through automated question generation with mutual supervision between models, enabling both generation and evaluation. The evaluation adopts a multiple-choice format, requiring models to select the correct answers from options containing multiple correct choices, thereby providing a more nuanced assessment of model confidence and judgment under uncertainty. Extensive experiments were conducted on 12 large language models, including ERNIE-Bot, ChatGLM, Qwen, and XVerse, with nine models exhibiting hallucination-free rates below 50%, underscoring the benchmark’s difficulty. Furthermore, HalluScope offers insights into hallucination-prone domains and hallucination types, providing guidance for fine-tuning models to mitigate hallucinations effectively.