<p>Speaker identification is a critical need in various applications but remains underexplored, especially for low-resource languages like Khmer, Javanese, Sudanese, and Nepali. This study introduces a novel hybrid model that combines MFCC and Chroma Short-Time Fourier Transform (Chroma STFT) features with a Transformer Encoder, targeting the unique challenges of speaker recognition in multilingual contexts. The paper demonstrates results using a carefully curated dataset with recordings from a diverse group of multilingual speakers. Our proposed model demonstrates excellent performance. An impressive accuracy of 99.04% is achieved. We obtain specifically for Khmer an overall accuracy of 98.16% with 98.19% for precision, 97.91% for recall, across all target languages with an F1 score of 97.95%. Chroma STFT is integrated with MFCC, with our approach, we also increase the ability to separate speakers, and add Transformer Encoder in multilingual and low resource languages like Sudanese, Nepali, Khmer, and Javanese. The design and implementation of speaker’s voice found in this research is further improved with valuable insight. Then they contribute significantly to the advancement of identification systems designed for multilingual settings. </p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Enhancing speaker identification in low-resource multilingual languages using hybrid MFCC-Chroma STFT and transformer encoder

  • Vajratiya Vajrobol,
  • Nitisha Aggarwal,
  • Geetika Jain Saxena,
  • Sanjeev Singh,
  • Amit Pundir

摘要

Speaker identification is a critical need in various applications but remains underexplored, especially for low-resource languages like Khmer, Javanese, Sudanese, and Nepali. This study introduces a novel hybrid model that combines MFCC and Chroma Short-Time Fourier Transform (Chroma STFT) features with a Transformer Encoder, targeting the unique challenges of speaker recognition in multilingual contexts. The paper demonstrates results using a carefully curated dataset with recordings from a diverse group of multilingual speakers. Our proposed model demonstrates excellent performance. An impressive accuracy of 99.04% is achieved. We obtain specifically for Khmer an overall accuracy of 98.16% with 98.19% for precision, 97.91% for recall, across all target languages with an F1 score of 97.95%. Chroma STFT is integrated with MFCC, with our approach, we also increase the ability to separate speakers, and add Transformer Encoder in multilingual and low resource languages like Sudanese, Nepali, Khmer, and Javanese. The design and implementation of speaker’s voice found in this research is further improved with valuable insight. Then they contribute significantly to the advancement of identification systems designed for multilingual settings.