Automatic Speaker Recognition Using Hybrid Parameters Based on Machine Learning Applied on Two Dataset
摘要
Automatic Speaker Recognition (ASR) is indeed crucial for enriching transcripts by identifying individual speakers from audio files. ASR algorithms extract unique speech characteristics from audio signals, providing valuable information about speakers within a document. These algorithms rely on feature extraction techniques that should be robust to recording parameters while containing sufficient speaker identity information. Machine learning plays a vital role in ASR, encompassing identification, verification, and detection of speakers. In a recent project, various machine learning algorithms such as k-nearest neighbors, SVM_SMO, MLP, Logistic function, and Naïve Bayes were applied to ASR, specifically used i-vectors for speaker recognition. These algorithms were evaluated on speaker identification and verification tasks across two different DataSet in English and Arabic languages. The evaluation criteria included the recognition rate (ASR Score/Accuracy), and the results highlighted the effectiveness of feature fusion for enhancing automatic speaker recognition performance.