Paraphrasing, the art of rephrasing text while retaining its original meaning, lies at the core of natural language understanding and generation. With the rise of demand for more domain-specialized models, high-quality data is more valued than ever; this includes paraphrasing. ParaFusion-Extend (PFE) is a large-scale dataset driven by Large Language Models incorporating lexical and phrasal knowledge. The dataset is curated to contain high-quality diverse paraphrase pairs and also separate knowledge bases that could be used for research work and data augmentation models. We show that PFE offers around at least a 30% increase in syntactic and lexical diversity compared to the original data sources that are commonly used. We demonstrate the effectiveness of PFE on several downstream tasks such as few-shot learning and training on sentence embeddings. We utilize a gold-standard evaluation scheme, which is further strengthened by human evaluation that shows the potential of PFE in advancing paraphrase generation.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

ParaFusion-Extended: Large Scale Paraphrase Dataset Integrating Lexico-Phrasal Knowledge

  • Lasal Jayawardena,
  • Prasan Yapa

摘要

Paraphrasing, the art of rephrasing text while retaining its original meaning, lies at the core of natural language understanding and generation. With the rise of demand for more domain-specialized models, high-quality data is more valued than ever; this includes paraphrasing. ParaFusion-Extend (PFE) is a large-scale dataset driven by Large Language Models incorporating lexical and phrasal knowledge. The dataset is curated to contain high-quality diverse paraphrase pairs and also separate knowledge bases that could be used for research work and data augmentation models. We show that PFE offers around at least a 30% increase in syntactic and lexical diversity compared to the original data sources that are commonly used. We demonstrate the effectiveness of PFE on several downstream tasks such as few-shot learning and training on sentence embeddings. We utilize a gold-standard evaluation scheme, which is further strengthened by human evaluation that shows the potential of PFE in advancing paraphrase generation.