Enhancing speaker identification in low-resource multilingual languages using hybrid MFCC-Chroma STFT and transformer encoder
摘要
Speaker identification is a critical need in various applications but remains underexplored, especially for low-resource languages like Khmer, Javanese, Sudanese, and Nepali. This study introduces a novel hybrid model that combines MFCC and Chroma Short-Time Fourier Transform (Chroma STFT) features with a Transformer Encoder, targeting the unique challenges of speaker recognition in multilingual contexts. The paper demonstrates results using a carefully curated dataset with recordings from a diverse group of multilingual speakers. Our proposed model demonstrates excellent performance. An impressive accuracy of 99.04% is achieved. We obtain specifically for Khmer an overall accuracy of 98.16% with 98.19% for precision, 97.91% for recall, across all target languages with an F1 score of 97.95%. Chroma STFT is integrated with MFCC, with our approach, we also increase the ability to separate speakers, and add Transformer Encoder in multilingual and low resource languages like Sudanese, Nepali, Khmer, and Javanese. The design and implementation of speaker’s voice found in this research is further improved with valuable insight. Then they contribute significantly to the advancement of identification systems designed for multilingual settings.