Text-Independent Voiceprints Identification Using Feed-Forward Back-Propagation with Layered Strategies
摘要
The task of identifying speakers forms a formidable field in Speaker Recognition which compares and matches the set of utterances spoken by an unknown or a known speaker with the available trained database of reference templates. And Deep Learning has proven to deliver better prediction and computational efficiency in many application domains of Speaker Recognition. Deep Neural Networks provide the best results as a backbone of deep learning due to their portability, versatility, and computational efficiency. This work presents a Fast Fourier Transform (FFT) based extraction of MFCC and GFCC features from a set of text-independent audios from the different number of speakers containing high ground noise. A Feed-Forward Back-propagation Neural Network (FFBNN) is used to categorize the voices of the selected speakers into tensor labels during the learning phase. These tensor labels are further tested with sample voice sets to identify the correct speaker. A comparative predictive outcome showed that FFBNN worked efficiently for 50 speakers, generating 85% of accuracy with 64–32 combinations of neurons at the hidden layer. GFCC features, designed to emulate the physiological characteristics of the human ear and mitigate the impact of noise, exhibit superior accuracy compared to MFCC features in the context of FFBNN.