Speech recognition systems are becoming popular because nowadays it is easy to capture audio signal through mobile devices. Moreover, biometric recognition systems are also popular because fingerprints, voice, retina, and face are specific to individuals and cannot be stolen easily as passwords and PINs. There are many datasets for speech recognition tasks. This research focuses on the benchmark for speaker verification SRE-08. There are different approaches for the automatic speaker verification task: prediction, regression, and classification. This research focuses only on classification with Probabilistic Neural Networks (PNN) and Emphasized Channel Attention, Propagation, and Aggregation in Time-Delay Neural Network (ECAPA-TDNN) methods. We use the Equal Error Rate (ERR) metric to measure the performance, and we found that the best performance corresponds to the classification approach (PNN) with the average EER = 13.3% (Accuracy 88%). The classification approach (PNN) achieves the single best with EER=0% (Accuracy 100%) on five speakers out of nine.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Artificial Neural Networks for Speaker Verification

  • Jose Luis Medellin-Garibay,
  • Juan C. Cuevas-Tello

摘要

Speech recognition systems are becoming popular because nowadays it is easy to capture audio signal through mobile devices. Moreover, biometric recognition systems are also popular because fingerprints, voice, retina, and face are specific to individuals and cannot be stolen easily as passwords and PINs. There are many datasets for speech recognition tasks. This research focuses on the benchmark for speaker verification SRE-08. There are different approaches for the automatic speaker verification task: prediction, regression, and classification. This research focuses only on classification with Probabilistic Neural Networks (PNN) and Emphasized Channel Attention, Propagation, and Aggregation in Time-Delay Neural Network (ECAPA-TDNN) methods. We use the Equal Error Rate (ERR) metric to measure the performance, and we found that the best performance corresponds to the classification approach (PNN) with the average EER = 13.3% (Accuracy 88%). The classification approach (PNN) achieves the single best with EER=0% (Accuracy 100%) on five speakers out of nine.