This paper explores the application of deep learning techniques for speaker identification in the under-resourced Amharic language, utilizing spectrogram images derived from the Amharic Speech Emotional Dataset (ASED). While traditional machine learning methods have been applied to speaker identification in Amharic, deep learning has yet to be fully explored for this task. We implement and evaluate four models: a custom Convolutional Neural Network (CNN) and three pretrained models (ResNet50, DenseNet121, and InceptionResNetV2), focusing on their ability to identify speakers from spectrogram images. The methodology involves converting audio recordings into spectrogram images, preprocessing them, and applying data augmentation to enhance generalization and reduce overfitting. Each model’s performance is assessed with and without augmentation, providing insights into the benefits and limitations of both approaches. Our results demonstrate that while the pretrained ResNet50 model achieves the highest accuracy (94.74%) on non-augmented data, the custom CNN benefits significantly from augmentation which helped it to manage overfitting more effectively compared to the pretrained models. The study highlights the potential of data augmentation for simpler models and the effectiveness of deep learning, offering a path for further research in Amharic speaker identification.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Speaker Identification in Amharic Language Using Spectrogram Images: Exploring Deep Learning Techniques

  • Bantegize Addis,
  • Birhaneselasie Abebe,
  • Million Meshesha

摘要

This paper explores the application of deep learning techniques for speaker identification in the under-resourced Amharic language, utilizing spectrogram images derived from the Amharic Speech Emotional Dataset (ASED). While traditional machine learning methods have been applied to speaker identification in Amharic, deep learning has yet to be fully explored for this task. We implement and evaluate four models: a custom Convolutional Neural Network (CNN) and three pretrained models (ResNet50, DenseNet121, and InceptionResNetV2), focusing on their ability to identify speakers from spectrogram images. The methodology involves converting audio recordings into spectrogram images, preprocessing them, and applying data augmentation to enhance generalization and reduce overfitting. Each model’s performance is assessed with and without augmentation, providing insights into the benefits and limitations of both approaches. Our results demonstrate that while the pretrained ResNet50 model achieves the highest accuracy (94.74%) on non-augmented data, the custom CNN benefits significantly from augmentation which helped it to manage overfitting more effectively compared to the pretrained models. The study highlights the potential of data augmentation for simpler models and the effectiveness of deep learning, offering a path for further research in Amharic speaker identification.