Clustering-based pseudo-labels (PLs) are widely used to optimize speaker embedding (SE) networks and train self-supervised speaker verification (SV) systems. However, PL-based self-supervised training depends on high-quality PLs, and clustering performance relies heavily on time- and resource-consuming data augmentation regularization. In this chapter, we propose an efficient and general-purpose multi-objective clustering algorithm that outperforms all other baselines for clustering SEs. Our approach, called Contrastive Information Maximization Clustering (CIMC), avoids explicit data augmentation, enabling fast training with low memory and compute resource usage. It is based on three principles: (1) self-augmented training to enforce representation invariance and maximize the information-theoretic dependency between samples and their predicted PLs; (2) virtual mixup training to impose local-Lipschitzness and enforce the cluster assumption; and (3) supervised contrastive learning to learn more discriminative features, pulling samples of the same class together and pushing samples of different classes apart, while improving robustness to natural corruptions. We provide a thorough comparative analysis of the performance of our clustering method against baselines using a variety of clustering metrics, demonstrating that we outperform all other clustering benchmarks. Moreover, we perform an ablation study to analyze the contribution of each component, including two other augmentation-based objectives, and show that our multi-objective approach provides beneficial complementary information. Finally, using the generated PLs to train our SE system enables us to achieve SOTA SV performance.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

An Efficient Clustering Algorithm for Self-Supervised Speaker Recognition

  • Abderrahim Fathan,
  • Jahangir Alam

摘要

Clustering-based pseudo-labels (PLs) are widely used to optimize speaker embedding (SE) networks and train self-supervised speaker verification (SV) systems. However, PL-based self-supervised training depends on high-quality PLs, and clustering performance relies heavily on time- and resource-consuming data augmentation regularization. In this chapter, we propose an efficient and general-purpose multi-objective clustering algorithm that outperforms all other baselines for clustering SEs. Our approach, called Contrastive Information Maximization Clustering (CIMC), avoids explicit data augmentation, enabling fast training with low memory and compute resource usage. It is based on three principles: (1) self-augmented training to enforce representation invariance and maximize the information-theoretic dependency between samples and their predicted PLs; (2) virtual mixup training to impose local-Lipschitzness and enforce the cluster assumption; and (3) supervised contrastive learning to learn more discriminative features, pulling samples of the same class together and pushing samples of different classes apart, while improving robustness to natural corruptions. We provide a thorough comparative analysis of the performance of our clustering method against baselines using a variety of clustering metrics, demonstrating that we outperform all other clustering benchmarks. Moreover, we perform an ablation study to analyze the contribution of each component, including two other augmentation-based objectives, and show that our multi-objective approach provides beneficial complementary information. Finally, using the generated PLs to train our SE system enables us to achieve SOTA SV performance.