Clarification need prediction (CNP) is a key task in conversational search, aiming to predict whether to ask a clarifying question or give an answer to the current user query. However, current research on CNP suffers from the issues of limited CNP training data and low efficiency. In this paper, we propose a zero-shot and efficient CNP framework (OUR), in which we first prompt LLMs in a zero-shot manner to generate two sets of synthetic queries: ambiguous and specific (unambiguous) queries. We then use the generated queries to train efficient CNP models. OUR eliminates the need for human-annotated clarification-need labels during training and avoids the use of LLMs with high query latency at query time. To further improve the generation quality of synthetic queries, we devise a topic-, information-need-, and query-aware CoT prompting strategy (PROMPT). Moreover, we enhance PROMPT with counterfactual query generation (SEQ), which guides LLMs first to generate a specific/ambiguous query and then sequentially generate its corresponding ambiguous/specific query. Experimental results show that OUR achieves superior CNP effectiveness and efficiency compared with zero- and few-shot LLM-based CNP predictors.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Zero-Shot and Efficient Clarification Need Prediction in Conversational Search

  • Lili Lu,
  • Chuan Meng,
  • Federico Ravenda,
  • Mohammad Aliannejadi,
  • Fabio Crestani

摘要

Clarification need prediction (CNP) is a key task in conversational search, aiming to predict whether to ask a clarifying question or give an answer to the current user query. However, current research on CNP suffers from the issues of limited CNP training data and low efficiency. In this paper, we propose a zero-shot and efficient CNP framework (OUR), in which we first prompt LLMs in a zero-shot manner to generate two sets of synthetic queries: ambiguous and specific (unambiguous) queries. We then use the generated queries to train efficient CNP models. OUR eliminates the need for human-annotated clarification-need labels during training and avoids the use of LLMs with high query latency at query time. To further improve the generation quality of synthetic queries, we devise a topic-, information-need-, and query-aware CoT prompting strategy (PROMPT). Moreover, we enhance PROMPT with counterfactual query generation (SEQ), which guides LLMs first to generate a specific/ambiguous query and then sequentially generate its corresponding ambiguous/specific query. Experimental results show that OUR achieves superior CNP effectiveness and efficiency compared with zero- and few-shot LLM-based CNP predictors.