<p>Speaker recognition (SR) refers to the means of personification through voice distinctiveness and it has been studied over many decades. Recently, owing to new advancements, SR has become a popular domain for study. The paper introduces the Self-Attention Augmented Wasserstein Generative Adversarial Network (SAA-WGAN) model along with the Hybrid Frilled Lizard Humboldt Squid Optimizer to be used on SR. It aims to recognize a speaker by extracting audio signals of voiceprints from a public dataset, followed by a model called Multi-Layer Adaptive Guided Side Window Box Filtering, applied for noise elimination in audio samples. After which, relevant features are extracted from the input signal using Fast Discrete Curvelet Transform (FDCT). In addition, the Fennec Fox optimization (FFO) model selects the most beneficial features. Then, the selected features are used by SAA-WGAN to perform SR based on speaker ID identification corresponding to each input voiceprint. To enhance the weight parameters of the SAA-WGAN model, a Hybrid Frilled Lizard Humboldt Squid Optimization Model (Hyb-FL-HSO) is proposed, combining optimization models such as FLO and HSOA. The suggested technique is implemented in Python. The strategy’s efficacy is comprehensively evaluated using evaluation metrics like accuracy, Matthew’s correlation coefficient (MCC), recall, kappa coefficient (KC), positive predictive value (PPV), and computation time (CT) and it is compared with other conventional methods. The overall accuracy of 98.9%, MCC of 96.9%, recall of 98.5%, and CT of 190.21s on enhancing the performance of SR.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Voiceprint revolution: a self-attention augmented Wasserstein generative adversarial network with hybrid frilled Lizard Humboldt squid optimization framework for speaker recognition

  • T. Rajesh Kumar,
  • C. Karthikeyan,
  • E. Rajesh Kumar,
  • N. Nirmal Singh

摘要

Speaker recognition (SR) refers to the means of personification through voice distinctiveness and it has been studied over many decades. Recently, owing to new advancements, SR has become a popular domain for study. The paper introduces the Self-Attention Augmented Wasserstein Generative Adversarial Network (SAA-WGAN) model along with the Hybrid Frilled Lizard Humboldt Squid Optimizer to be used on SR. It aims to recognize a speaker by extracting audio signals of voiceprints from a public dataset, followed by a model called Multi-Layer Adaptive Guided Side Window Box Filtering, applied for noise elimination in audio samples. After which, relevant features are extracted from the input signal using Fast Discrete Curvelet Transform (FDCT). In addition, the Fennec Fox optimization (FFO) model selects the most beneficial features. Then, the selected features are used by SAA-WGAN to perform SR based on speaker ID identification corresponding to each input voiceprint. To enhance the weight parameters of the SAA-WGAN model, a Hybrid Frilled Lizard Humboldt Squid Optimization Model (Hyb-FL-HSO) is proposed, combining optimization models such as FLO and HSOA. The suggested technique is implemented in Python. The strategy’s efficacy is comprehensively evaluated using evaluation metrics like accuracy, Matthew’s correlation coefficient (MCC), recall, kappa coefficient (KC), positive predictive value (PPV), and computation time (CT) and it is compared with other conventional methods. The overall accuracy of 98.9%, MCC of 96.9%, recall of 98.5%, and CT of 190.21s on enhancing the performance of SR.