Current research on dialogue contradiction detection primarily focuses on single-hop reasoning. We introduce Multi-hop ContraDial, a dataset specifically designed to address contradictions requiring multi-hop reasoning, constructed through an automated process that extends existing multi-premise natural language inference datasets. The dataset consists of 8,028 samples, including both contradictory and non-contradictory dialogue instances. Experimental results demonstrate that models fine-tuned on Multi-hop ContraDial outperform those trained on single-hop contradiction datasets, underscoring the necessity of specialized data for handling multi-hop reasoning. Notably, GPT-4o achieves an MCC of only 0.5561 on the Sub2 subset, highlighting the challenge of complex logical dependencies. In contrast, our task-adapted model reaches 0.8892, demonstrating the effectiveness of domain-specific fine-tuning. This work provides an important resource for advancing dialogue systems capable of maintaining consistency in complex reasoning scenarios.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Multi-hop ContraDial: Multi-hop Reasoning for Contradiction Detection in Dialogue

  • Tao Lin,
  • Dongning Rao

摘要

Current research on dialogue contradiction detection primarily focuses on single-hop reasoning. We introduce Multi-hop ContraDial, a dataset specifically designed to address contradictions requiring multi-hop reasoning, constructed through an automated process that extends existing multi-premise natural language inference datasets. The dataset consists of 8,028 samples, including both contradictory and non-contradictory dialogue instances. Experimental results demonstrate that models fine-tuned on Multi-hop ContraDial outperform those trained on single-hop contradiction datasets, underscoring the necessity of specialized data for handling multi-hop reasoning. Notably, GPT-4o achieves an MCC of only 0.5561 on the Sub2 subset, highlighting the challenge of complex logical dependencies. In contrast, our task-adapted model reaches 0.8892, demonstrating the effectiveness of domain-specific fine-tuning. This work provides an important resource for advancing dialogue systems capable of maintaining consistency in complex reasoning scenarios.