Multi-round Q&A Hallucination Analysis of LLMs in Geographical Education
摘要
Large language models (LLMs) have become increasingly important in education, for their potential in personalized learning and interactive tutoring. However, multi-round educational Q&A may trigger hallucinations, which can harm teaching credibility and pose risks. Notably, current research lacks evaluation datasets to detect and mitigate hallucinations in educational contexts. To address this issue, we introduce the first multi-round Q&A hallucination detection dataset tailored for geographical education. This dataset employs a two-stage approach to simulate authentic educational Q&A and achieves precise hallucination injection. Firstly, we manually build standardized multi-round Q&A records using real geographical education questions. In the second stage, we design a systematic framework integrating 8 hallucination patterns for reliable and high-quality automated generation. Furthermore, we develop a method that combines chain-of-thought distillation with MoBA guidance module to alleviate specific hallucinations in LLMs, like implicit premise loss. We evaluate six popular LLMs with comprehensive metrics and analyze hallucination distributions. The results indicate LLMs are prone to generating hallucinations under this scenario, highlighting challenges to their knowledge and logical capabilities. Comparative experiments also validate that our proposed method help alleviate hallucination issues of LLMs in multi-round educational Q&A.