Layer-wise symbolic attention instability as a diagnostic signal for hallucination in large language models
摘要
Large Language Models (LLMs) tend to hallucinate when processing symbolically complex linguistic structures. Existing literature evaluates hallucination either through their mechanistic interpretability or at the behavioral output level, but hardly links the symbolic triggers to their layer-wise representational causes. This paper introduces a unified symbolic, behavioral, and mechanistic framework that connects symbolic triggers with internal failure dynamics in transformer architectures. The study evaluates five open-weight LLMs across QA, MCQ, and Odd-One-Out formats on the HaluEval and TruthfulQA datasets, focusing on negation, exceptions, modifiers, numbers, and named entity cues. The results show that hallucination rates remain high across all model scales, with all symbolic categories exhibiting high hallucination rates, and exceptions and numbers often showing comparable or higher values across models. Constrained task formats reduce surface errors but preserve failure patterns, indicating representational instability rather than purely decoding artifacts. Layer-wise analysis shows peak symbolic attention variance in early transformer layers (2–4), after which these patterns persist across deep layers. The consistency of this behavior across architectures suggests that hallucination is strongly associated with weakness in symbolic encoding. The framework provides an interpretable basis for diagnosing and stabilizing symbolic reasoning in LLMs.