A systematic review of accent classification techniques and datasets for inclusive speech recognition
摘要
Automatic speech recognition (ASR) systems have made significant progress in recent years, yet they continue to struggle with accented and low-resource speech, particularly from underrepresented linguistic communities. This systematic review explores advances in accent classification and accent similarity modelling, focusing on machine learning (ML) and deep learning (DL) methods. We analyse 58 peer-reviewed studies and 24 publicly available speech datasets, categorising the latter based on recording hours, sampling rates, and speaker demographics. The review identifies core approaches, including convolutional neural networks (CNNs), long short-term memory (LSTM) networks, Transformer-based models, and multi-embedding strategies that improve accent-aware ASR performance. Despite these advances, critical challenges persist, including limited data availability for non-native accents, a lack of standardised evaluation metrics, and minimal exploration of geographically proximate accent variations in low-resource settings. To address these gaps, we highlight future research directions involving pronunciation distance metrics, accent-aware model adaptation, large language models (LLMs), and domain generalisation techniques. This review provides a valuable foundation for the development of more inclusive and robust ASR systems capable of supporting linguistic diversity in real-world applications.