<p>Contrastive learning is an effective method of self-supervised representation learning, which has made significant strides in the application of sentence representation learning in recent years. However, most previous studies typically focus on generating effective positive samples through various data augmentation methods, overlooking the fact that negative samples are often randomly sampled from the training data. This practice may introduce false negatives leading to negative sampling bias, which can consequently impair the performance of contrastive learning models. This issue has garnered considerable attention, yet current solutions have some limitations in dealing with negative sampling bias in a relatively straightforward way. To tackle these challenges, this paper proposes a <Emphasis Type="BoldUnderline">D</Emphasis>ifficulty-based <Emphasis Type="BoldUnderline">D</Emphasis>ebiased <Emphasis Type="BoldUnderline">C</Emphasis>ontrastive <Emphasis Type="BoldUnderline">L</Emphasis>earning (DDCL) framework to mitigate the effects of negative sampling bias. Specifically, our framework introduces a relative difficulty strategy to accurately screen out potential false negatives from the negative samples. Additionally, we propose a similarity-adaptive weighting method to alleviate negative sampling bias, which dynamically assigns different weights to the screened false negative samples. Furthermore, we incorporate the negation of anchor samples into contrastive learning, which can enable the model to distinguish between structural and semantic similarity of sentences. Experiments across seven semantic text similarity tasks show that our method outperforms baseline models, with an average performance improvement to 78.8% Spearman’s correlation.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Self-supervised contrastive learning of sentence representation with difficulty-based sampling

  • Jian Wu,
  • Dong Qiu

摘要

Contrastive learning is an effective method of self-supervised representation learning, which has made significant strides in the application of sentence representation learning in recent years. However, most previous studies typically focus on generating effective positive samples through various data augmentation methods, overlooking the fact that negative samples are often randomly sampled from the training data. This practice may introduce false negatives leading to negative sampling bias, which can consequently impair the performance of contrastive learning models. This issue has garnered considerable attention, yet current solutions have some limitations in dealing with negative sampling bias in a relatively straightforward way. To tackle these challenges, this paper proposes a Difficulty-based Debiased Contrastive Learning (DDCL) framework to mitigate the effects of negative sampling bias. Specifically, our framework introduces a relative difficulty strategy to accurately screen out potential false negatives from the negative samples. Additionally, we propose a similarity-adaptive weighting method to alleviate negative sampling bias, which dynamically assigns different weights to the screened false negative samples. Furthermore, we incorporate the negation of anchor samples into contrastive learning, which can enable the model to distinguish between structural and semantic similarity of sentences. Experiments across seven semantic text similarity tasks show that our method outperforms baseline models, with an average performance improvement to 78.8% Spearman’s correlation.