SP-Aug: Towards Efficient Semantic-Preserving Augmentations in Contrastive Learning via Hierarchical Outlier Factor
摘要
Data augmentation is an effective way to generate abundant synthetic data from original data, with the same labels preserved. It has recently demonstrated its significant advantages in contrastive learning by constructing positive pairs with augmented data. However, the implicit underlying assumption that data augmentation is consistently semantic-preserving is unrealistic and we observe that not all augmentations are beneficial for contrastive learning. Yet it is challenging to distinguish which augmentations preserve the semantic consistency of labels for a specific task. To tackle this challenge, we formulate the problem of selecting optimal augmentations as a multi-level anomaly detection problem, and propose a novel method towards efficient Semantic-Preserving Augmentations (SP-Aug) to leverage intrinsic hierarchical clustering information to guide the detection of unfavorable augmentations from local and intra-cluster perspectives. We design a new metric named Hierarchical Outlier Factor (HOF) to measure the multi-level semantic inconsistency of augmentations across different clustering partitions. Empirically we evaluate our method on 3 popular contrastive learning models and compare it with 3 state-of-the-art augmentation selection methods, demonstrating the superiority of ours in discovering outliers and identifying semantic-preserving augmentations in classification tasks.