<p>Prompt-based learning has shown considerable promise in reformulating various downstream tasks as cloze problems by combining original input with a predetermined template. This approach demonstrates its effectiveness, especially in few-shot learning scenarios, where the model is trained on a scarce amount of data. Despite its successes, the limited templates and text in few-shot prompt-based learning scenarios leave significant room for performance improvement. Moreover, existing methods sometimes resort to model ensembles, which, while effective, could potentially hamper model efficiency due to increased computational demands [<CitationRef CitationID="CR1">1</CitationRef>]. To address these issues, we introduce <span>MixPro</span>, an augmentation method designed to augment both the vanilla input text and the templates. We implement this through the token-level, the sentence-level, and the template-level Mixup strategies. We conduct experiments on five few-shot datasets, and the results show that our <span>MixPro</span> achieves an average performance improvement of 5.08% compared to the backbone model before augmentation. Moreover, it outperforms other augmentation baselines, demonstrating its superior effectiveness.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

MixPro: Simple yet Effective Data Augmentation for Prompt-based Learning

  • Bohan Li,
  • Longxu Dou,
  • Yutai Hou,
  • Yunlong Feng,
  • Honglin Mu,
  • Enbo Wang,
  • Qingfu Zhu,
  • Qinghua Sun,
  • Wanxiang Che

摘要

Prompt-based learning has shown considerable promise in reformulating various downstream tasks as cloze problems by combining original input with a predetermined template. This approach demonstrates its effectiveness, especially in few-shot learning scenarios, where the model is trained on a scarce amount of data. Despite its successes, the limited templates and text in few-shot prompt-based learning scenarios leave significant room for performance improvement. Moreover, existing methods sometimes resort to model ensembles, which, while effective, could potentially hamper model efficiency due to increased computational demands [1]. To address these issues, we introduce MixPro, an augmentation method designed to augment both the vanilla input text and the templates. We implement this through the token-level, the sentence-level, and the template-level Mixup strategies. We conduct experiments on five few-shot datasets, and the results show that our MixPro achieves an average performance improvement of 5.08% compared to the backbone model before augmentation. Moreover, it outperforms other augmentation baselines, demonstrating its superior effectiveness.