Text semantic matching constitutes a foundational challenge in the domain of natural language processing (NLP). Unsupervised text matching methodologies have gained substantial attention for their capacity to address the issue of missing data in downstream text matching datasets. In recent years, the introduction of deep pretrained language models like MacBERT, based on the bidirectional encoder representations from transformers (BERT) model, has provided improved solutions for many Chinese natural language processing tasks. This has significantly enhanced the performance of numerous downstream tasks, leading to better results. However, a salient issue that afflicts such models is their susceptibility to anisotropy, a phenomenon that impedes their full capacity to harness underlying semantic features. In this paper, we have implemented an unsupervised text semantic matching algorithm that relies on contrastive learning. We utilize contrastive learning techniques to fine-tune the MacBERT model, effectively mitigating the anisotropy issue in the model's output sentence vectors. To substantiate our claims, we conduct a series of experiments on three authentic datasets. The experimental results affirm the superior performance of the proposed method. Finally, we also examine the influence of various existing data augmentation methods in contrastive learning on the overall performance.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Unsupervised Text Semantic Matching Based on Contrastive Learning

  • Chao Chen,
  • Ronghui Zhang,
  • Chao Wei,
  • Xiaojun Jing

摘要

Text semantic matching constitutes a foundational challenge in the domain of natural language processing (NLP). Unsupervised text matching methodologies have gained substantial attention for their capacity to address the issue of missing data in downstream text matching datasets. In recent years, the introduction of deep pretrained language models like MacBERT, based on the bidirectional encoder representations from transformers (BERT) model, has provided improved solutions for many Chinese natural language processing tasks. This has significantly enhanced the performance of numerous downstream tasks, leading to better results. However, a salient issue that afflicts such models is their susceptibility to anisotropy, a phenomenon that impedes their full capacity to harness underlying semantic features. In this paper, we have implemented an unsupervised text semantic matching algorithm that relies on contrastive learning. We utilize contrastive learning techniques to fine-tune the MacBERT model, effectively mitigating the anisotropy issue in the model's output sentence vectors. To substantiate our claims, we conduct a series of experiments on three authentic datasets. The experimental results affirm the superior performance of the proposed method. Finally, we also examine the influence of various existing data augmentation methods in contrastive learning on the overall performance.