This research investigates the development approach to developing a speech recognition model for Thai Lanna vocabulary, explicitly addressing the challenges of confused vocabulary items in the Kham Mueang dialect. The research aims to create a robust system capable of accurately recognizing and interpreting the nuances of Northern Thai speech by utilizing deep learning techniques, particularly speech transformer models. The research is based on a unique Confused Thai Lanna Vocabulary dataset, meticulously collected between 2023 and 2024. It comprises authentic recordings from native speakers and fluent communicators. The speech recognition models - HuBERT, Wav2Vec2-TH, Wav2Vec2, and WavLM - were employed to enhance the recognition and interpretation of the Confused Thai Lanna Vocabulary. Results indicate that all models effectively adapt to the Lanna language, with HuBERT consistently outperforming across all metrics. Wav2Vec2-TH demonstrated the most balanced performance across different word categories, while WavLM excelled in recognizing training words. However, significant challenges remain in processing new words and accent variants, highlighting areas for future research. This study advances ASR technology for regional dialects and contributes to preserving linguistic diversity and promoting inclusive communication in Northern Thailand.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Speech Recognition Model for Confused Thai Lanna Vocabulary Using Deep Learning Techniques

  • Wongpanya S. Nuankaew,
  • Pathapol Jomsawan,
  • Thapanapong Sararat,
  • Pratya Nuankaew

摘要

This research investigates the development approach to developing a speech recognition model for Thai Lanna vocabulary, explicitly addressing the challenges of confused vocabulary items in the Kham Mueang dialect. The research aims to create a robust system capable of accurately recognizing and interpreting the nuances of Northern Thai speech by utilizing deep learning techniques, particularly speech transformer models. The research is based on a unique Confused Thai Lanna Vocabulary dataset, meticulously collected between 2023 and 2024. It comprises authentic recordings from native speakers and fluent communicators. The speech recognition models - HuBERT, Wav2Vec2-TH, Wav2Vec2, and WavLM - were employed to enhance the recognition and interpretation of the Confused Thai Lanna Vocabulary. Results indicate that all models effectively adapt to the Lanna language, with HuBERT consistently outperforming across all metrics. Wav2Vec2-TH demonstrated the most balanced performance across different word categories, while WavLM excelled in recognizing training words. However, significant challenges remain in processing new words and accent variants, highlighting areas for future research. This study advances ASR technology for regional dialects and contributes to preserving linguistic diversity and promoting inclusive communication in Northern Thailand.