Task-oriented chatbots rely on extensive and diverse dataset of user utterances to recognize tasks and intents effectively. Traditionally, these datasets are created through crowdsourcing, where crowd workers expand an initial set of seed utterances through paraphrasing. Although effective, crowdsourcing techniques have disadvantages, including high cost, time-consuming, low output quality, and often focusing only on lexical diversity. The emergence of Large Language Models (LLMs) presents a promising alternative for generating high-quality and diverse paraphrases more efficiently. In this paper, we investigate whether LLM can replace the crowd in generating syntactically diverse paraphrases to train chatbots. We replicate an existing crowdsourcing workflow using GPT, maintaining similar scale, prompts, and data. We evaluate the data across three dimensions: (1) comparing paraphrases generated by crowd workers and GPT; (2) examining the impact of changing the number of paraphrases per request; and (3) assessing performance across different prompt strategies. Our findings reveal that GPT generated paraphrases with greater syntactical diversity and semantic relevance while resulting in a 98% cost reduction compared with crowdsourcing.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

LLMs to Replace Crowdsourcing in Generating Syntactically Diverse Paraphrases for Task-Oriented Chatbots

  • Auday Berro,
  • Vitor Gaboardi dos Santos,
  • Boualem Benatallah,
  • Khalid Benabdeslem

摘要

Task-oriented chatbots rely on extensive and diverse dataset of user utterances to recognize tasks and intents effectively. Traditionally, these datasets are created through crowdsourcing, where crowd workers expand an initial set of seed utterances through paraphrasing. Although effective, crowdsourcing techniques have disadvantages, including high cost, time-consuming, low output quality, and often focusing only on lexical diversity. The emergence of Large Language Models (LLMs) presents a promising alternative for generating high-quality and diverse paraphrases more efficiently. In this paper, we investigate whether LLM can replace the crowd in generating syntactically diverse paraphrases to train chatbots. We replicate an existing crowdsourcing workflow using GPT, maintaining similar scale, prompts, and data. We evaluate the data across three dimensions: (1) comparing paraphrases generated by crowd workers and GPT; (2) examining the impact of changing the number of paraphrases per request; and (3) assessing performance across different prompt strategies. Our findings reveal that GPT generated paraphrases with greater syntactical diversity and semantic relevance while resulting in a 98% cost reduction compared with crowdsourcing.