Beyond Factualism: A Study of LLM Calibration Through the Lens of Conversational Emotion Recognition
摘要
The calibration of large language models (LLMs) is crucial for ensuring their reliability and effectiveness, especially in tasks requiring sophisticated reasoning. While significant research has focused on understanding the calibration of LLMs in factual tasks like question answering, there is a notable gap in assessing their calibration for more complex tasks. This gap is particularly evident in domains such as emotion recognition in conversation (ERC), where the challenge extends beyond basic contextual understanding to interpreting nuanced emotional states integral to human interactions. In this paper, we explore the extent to which state-of-the-art LLMs are well-calibrated for the specific task of ERC. Our findings reveal that these models exhibit poor calibration—specifically, a tendency toward overconfidence—such that their confidence levels do not accurately reflect actual performance. While one can leverage the intrinsic verification capabilities of LLMs to assess the correctness of their predictions to some degree, this does not sufficiently address the significant issue of overconfidence inherent in these models.