In the field of computer vision, Self-Supervised Learning (SSL) is an emerging AI paradigm that leverages unlabeled data to generate supervisory signals, allowing the model to learn useful feature representations without requiring manual labeling. SSL generates labels from unlabeled data by designing pretext tasks to learn features. Once the model acquires these useful feature representations, they can be transferred to downstream tasks. However, studies have shown that this transfer process is vulnerable to backdoor attacks. We propose a model-based backdoor attack approach named TNSSL (TrojanNet for Self-Supervised Learning), which is the first neural network Trojan attack targeting the SSL domain. In our approach, we inject a trained neural network Trojan into the pre-trained encoder, distribute the Trojan via third-party platforms, and activate the Trojan during the construction of specific classifiers for downstream tasks. Instead of modifying the parameters of the downstream classifier, we insert a small Trojan module. Experimental results demonstrate TNSSL’s effectiveness in the following aspects: (1) It is triggered by a small signal and remains unaffected by other noise; (2) It can inject multiple triggers simultaneously, enabling multi-target backdoor attacks with a 100% success rate, without impacting the original task; (3) Compared to traditional backdoor attacks targeting SSL, its training-free mechanism significantly reduces the workload required for encoder training.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

TNSSL: TrojanNet Attack in Self-supervised Learning

  • Wenyin Yang,
  • Guihong Sun,
  • Li Ma,
  • Zikai Zhao,
  • Xianxi Liang,
  • Yali Ma

摘要

In the field of computer vision, Self-Supervised Learning (SSL) is an emerging AI paradigm that leverages unlabeled data to generate supervisory signals, allowing the model to learn useful feature representations without requiring manual labeling. SSL generates labels from unlabeled data by designing pretext tasks to learn features. Once the model acquires these useful feature representations, they can be transferred to downstream tasks. However, studies have shown that this transfer process is vulnerable to backdoor attacks. We propose a model-based backdoor attack approach named TNSSL (TrojanNet for Self-Supervised Learning), which is the first neural network Trojan attack targeting the SSL domain. In our approach, we inject a trained neural network Trojan into the pre-trained encoder, distribute the Trojan via third-party platforms, and activate the Trojan during the construction of specific classifiers for downstream tasks. Instead of modifying the parameters of the downstream classifier, we insert a small Trojan module. Experimental results demonstrate TNSSL’s effectiveness in the following aspects: (1) It is triggered by a small signal and remains unaffected by other noise; (2) It can inject multiple triggers simultaneously, enabling multi-target backdoor attacks with a 100% success rate, without impacting the original task; (3) Compared to traditional backdoor attacks targeting SSL, its training-free mechanism significantly reduces the workload required for encoder training.