Traditional Dialogue State Tracking (DST) evaluation relies heavily on large amounts of labeled data and employs exact matching, without incorporating any language understanding ability, often resulting in inaccurate assessments. Recently, leveraging large language models (LLMs) in evaluating natural language processing tasks has shown promising results. However, using LLMs for DST evaluation remains underexplored. In this paper, we propose an Error-Guided Learning for multidimensional evaluation method to improve the performance of zero-shot DST evaluation. Specifically, it begins with iterative learning from existing DST evaluation error cases to automatically generate distinct error prompts using an LLM. These categorized error prompts are then combined with a prompt template to evaluate the DST model across three dimensions, reducing misassessment and providing detailed explanations. Finally, the verifier consolidates the assessment results from each dimension to reach a final decision, utilizing a double-checking mechanism to mitigate LLM hallucination. Experimental results demonstrate that our method achieves state-of-the-art (SOTA) performance compared to baselines, establishing a robust foundation for more comprehensive and interpretable zero-shot DST evaluation( \(^4\) https://github.com/Kurak0oo/DST_EVAL ).

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

EGL-DST: Error-Guided Learning for Multidimensional Evaluation Method of Dialogue State Tracking via GPT-4

  • Wenjie Dong,
  • Sirong Chen,
  • Ming Gu,
  • Yan Yang

摘要

Traditional Dialogue State Tracking (DST) evaluation relies heavily on large amounts of labeled data and employs exact matching, without incorporating any language understanding ability, often resulting in inaccurate assessments. Recently, leveraging large language models (LLMs) in evaluating natural language processing tasks has shown promising results. However, using LLMs for DST evaluation remains underexplored. In this paper, we propose an Error-Guided Learning for multidimensional evaluation method to improve the performance of zero-shot DST evaluation. Specifically, it begins with iterative learning from existing DST evaluation error cases to automatically generate distinct error prompts using an LLM. These categorized error prompts are then combined with a prompt template to evaluate the DST model across three dimensions, reducing misassessment and providing detailed explanations. Finally, the verifier consolidates the assessment results from each dimension to reach a final decision, utilizing a double-checking mechanism to mitigate LLM hallucination. Experimental results demonstrate that our method achieves state-of-the-art (SOTA) performance compared to baselines, establishing a robust foundation for more comprehensive and interpretable zero-shot DST evaluation( \(^4\) https://github.com/Kurak0oo/DST_EVAL ).