<p>Automatic speech recognition (ASR) systems have made significant progress in recent years, yet they continue to struggle with accented and low-resource speech, particularly from underrepresented linguistic communities. This systematic review explores advances in accent classification and accent similarity modelling, focusing on machine learning (ML) and deep learning (DL) methods. We analyse 58 peer-reviewed studies and 24 publicly available speech datasets, categorising the latter based on recording hours, sampling rates, and speaker demographics. The review identifies core approaches, including convolutional neural networks (CNNs), long short-term memory (LSTM) networks, Transformer-based models, and multi-embedding strategies that improve accent-aware ASR performance. Despite these advances, critical challenges persist, including limited data availability for non-native accents, a lack of standardised evaluation metrics, and minimal exploration of geographically proximate accent variations in low-resource settings. To address these gaps, we highlight future research directions involving pronunciation distance metrics, accent-aware model adaptation, large language models (LLMs), and domain generalisation techniques. This review provides a valuable foundation for the development of more inclusive and robust ASR systems capable of supporting linguistic diversity in real-world applications.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A systematic review of accent classification techniques and datasets for inclusive speech recognition

  • Amina Salifu,
  • Henry Nunoo Mensah,
  • Eric Tutu Tchao,
  • Ali Musah Ibrahim,
  • Francisca Adoma Acheampong,
  • Jerry John Kponyo,
  • Andrew Selasi Agbemenu

摘要

Automatic speech recognition (ASR) systems have made significant progress in recent years, yet they continue to struggle with accented and low-resource speech, particularly from underrepresented linguistic communities. This systematic review explores advances in accent classification and accent similarity modelling, focusing on machine learning (ML) and deep learning (DL) methods. We analyse 58 peer-reviewed studies and 24 publicly available speech datasets, categorising the latter based on recording hours, sampling rates, and speaker demographics. The review identifies core approaches, including convolutional neural networks (CNNs), long short-term memory (LSTM) networks, Transformer-based models, and multi-embedding strategies that improve accent-aware ASR performance. Despite these advances, critical challenges persist, including limited data availability for non-native accents, a lack of standardised evaluation metrics, and minimal exploration of geographically proximate accent variations in low-resource settings. To address these gaps, we highlight future research directions involving pronunciation distance metrics, accent-aware model adaptation, large language models (LLMs), and domain generalisation techniques. This review provides a valuable foundation for the development of more inclusive and robust ASR systems capable of supporting linguistic diversity in real-world applications.