Evaluation of Large Language Models’ True Belief Inference Capability
摘要
Large language models (LLMs) have shown remarkable abilities in reasoning tasks, especially in Theory of Mind (ToM). However, most evaluations have focused on their performance on false beliefs tasks, ignoring correct beliefs. In this paper, we constructed the correct belief task dataset, improved the traditional method of using static data sets in the field of theory of mind, and tested it in the way of a dynamic gamified assessment (DGA). There are two kinds of content for the correct belief test, one is the constant belief test, and the other is the belief that will change. The experimental results show that most of the LLMs perform well in the field of correct beliefs, they need to revise their beliefs step by step with the assistance of another agent, and a small number of models are not good enough to understand the task content.