Speech is a fundamental aspect of human communication, driving advancements in Automatic Speech Recognition (ASR) to bridge the gap between humans and machines. ASR has evolved significantly with the introduction of Deep Neural Networks (DNNs), which enable models to learn speech patterns directly from raw audio. The transition to end-to-end DNNs architecture, including models like Whisper and Conformer, dramatically improved ASR performance. In the context of Thai ASR, researchers have adopted DNNs-based models, including fine-tuned versions of Whisper, to improve recognition accuracy. However, prior studies have shown that Thai ASR still faces challenges due to the lack of spaces in Thai sentences and regional dialectal variations. To address these challenges, this study proposes an alternative approach by modifying the existing Whisper model by integrating the Conformer architecture. We refer to the resulting model as Whisper-Conformer. The model is trained on 373.5 h of Thai speech data. Our results indicate that Whisper-Conformer learns significantly faster than baseline models and outperforms the fine-tuned Whisper model, achieving 0.64 Word Error Rate (WER) and 0.42 Character Error Rate (CER) on the Common Voice Corpus (v18) and 83.27 WER and 39.96 CER on the Thai Dialect Corpus, without using a language model for spelling correction. These findings suggest that integrating the Conformer architecture enhances ASR performance and enables the model to handle challenges in Thai ASR more effectively. The pretrained models are available at https://huggingface.co/Thanakron/whisperConformer-medium-th .

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Whisper-Conformer: A Modified Automatic Speech Recognition for Thai Speech Recognition

  • Thanakron Noppanamas,
  • Suronapee Phooomvuthisarn

摘要

Speech is a fundamental aspect of human communication, driving advancements in Automatic Speech Recognition (ASR) to bridge the gap between humans and machines. ASR has evolved significantly with the introduction of Deep Neural Networks (DNNs), which enable models to learn speech patterns directly from raw audio. The transition to end-to-end DNNs architecture, including models like Whisper and Conformer, dramatically improved ASR performance. In the context of Thai ASR, researchers have adopted DNNs-based models, including fine-tuned versions of Whisper, to improve recognition accuracy. However, prior studies have shown that Thai ASR still faces challenges due to the lack of spaces in Thai sentences and regional dialectal variations. To address these challenges, this study proposes an alternative approach by modifying the existing Whisper model by integrating the Conformer architecture. We refer to the resulting model as Whisper-Conformer. The model is trained on 373.5 h of Thai speech data. Our results indicate that Whisper-Conformer learns significantly faster than baseline models and outperforms the fine-tuned Whisper model, achieving 0.64 Word Error Rate (WER) and 0.42 Character Error Rate (CER) on the Common Voice Corpus (v18) and 83.27 WER and 39.96 CER on the Thai Dialect Corpus, without using a language model for spelling correction. These findings suggest that integrating the Conformer architecture enhances ASR performance and enables the model to handle challenges in Thai ASR more effectively. The pretrained models are available at https://huggingface.co/Thanakron/whisperConformer-medium-th .