Developing a Vietnamese Regional Voice Dataset and Benchmark for Region Recognition Based on Speech
摘要
To enhance human-computer interaction, enable personalized customer service, and improve accessibility for diverse populations across various regions, recognizing regional accents in speech is a critical challenge. In this paper, we present a comprehensive speech dataset comprising over 9,000 audio samples from 63 provinces in Vietnam, representing six distinct regional accents. To provide benchmarks and demonstrate the challenges posed by our dataset regarding dialectal accents, we apply various machine learning and deep learning techniques to the downstream task of regional accent identification. The experiments achieve classification accuracies ranging from 71.01% to 83.17%. These empirical results highlight both the influence of geographical factors on accent and the limitations of current approaches in speech recognition tasks involving multi-dialect speech data. Extracting regional speech features and designing robust classification algorithms remain challenging tasks. This dataset, along with the performance comparison of common models, provides valuable insights for developing more efficient speech classification techniques.