<p>Text augmentation, a technique for generating new samples through various combinations, noise, and manipulations of small datasets, is an essential technique in natural language processing research. This methodology enables the construction of robust models during the training step by enhancing data diversity. However, determining the manipulation level remains a significant challenge. When the manipulation intensity is too low, insufficient data diversity is generated, leading to suboptimal augmentation effects. Conversely, excessive manipulation can compromise label reliability, resulting in a degradation of model performance. To address the challenge of “manipulation level,” we propose a text augmentation technique that can make systematic adjustments. In particular, we introduce a method for flexibly resetting the range of the candidate pool for manipulations, ensuring an optimal level of randomness during the augmentation process. We also introduce an advanced sentence embedding that supports reliable pseudo-labeling across different manipulation levels. Additionally, we utilize ChatGPT model in the final stage to enhance the coherence and expressiveness of the generated text, thereby improving the quality of the output. To evaluate the effectiveness of our approach, we performed comparisons with existing text augmentation approaches. The experimental results show significant performance improvements in almost all test datasets.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Text augmentation method with adjustable manipulation intensity based on in-context learning

  • Yuho Cha,
  • Younghoon Lee

摘要

Text augmentation, a technique for generating new samples through various combinations, noise, and manipulations of small datasets, is an essential technique in natural language processing research. This methodology enables the construction of robust models during the training step by enhancing data diversity. However, determining the manipulation level remains a significant challenge. When the manipulation intensity is too low, insufficient data diversity is generated, leading to suboptimal augmentation effects. Conversely, excessive manipulation can compromise label reliability, resulting in a degradation of model performance. To address the challenge of “manipulation level,” we propose a text augmentation technique that can make systematic adjustments. In particular, we introduce a method for flexibly resetting the range of the candidate pool for manipulations, ensuring an optimal level of randomness during the augmentation process. We also introduce an advanced sentence embedding that supports reliable pseudo-labeling across different manipulation levels. Additionally, we utilize ChatGPT model in the final stage to enhance the coherence and expressiveness of the generated text, thereby improving the quality of the output. To evaluate the effectiveness of our approach, we performed comparisons with existing text augmentation approaches. The experimental results show significant performance improvements in almost all test datasets.