Speech recognition is a crucial technology that enables human–computer interaction through natural language processing. In the context of the Bengali language, developing an accurate and efficient speech recognition system remains challenging due to its phonetic complexities and limited available resources. This research presents a comprehensive study on enhancing Bengali speech recognition by leveraging spectrogram analysis and deep learning techniques. We first investigate the fundamental aspects of spectrogram analysis to extract meaningful acoustic features from the Bengali speech signals. Next, we explore the efficacy of deep learning (DL) models using convolutional neural networks (CNNs) in capturing intricate linguistic patterns from spectrogram representations. Our research involves the construction of a medium-scale Bengali speech dataset, facilitating rigorous evaluation of the proposed methods. The experimental outcomes illustrate the effectiveness of our technique, showcasing significant advancements in Bengali speech recognition accuracy and performance. The findings presented in this paper hold promising implications for various applications, including voice-controlled systems, language translation, and transcription services in the Bengali-speaking community.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Empowering Bangla Speech Recognition System Through Spectrogram Analysis and Deep Learning Approach

  • Bachchu Paul,
  • Saloni Sahal,
  • Sumita Guchhait,
  • Santanu Manna,
  • Utpal Nandi

摘要

Speech recognition is a crucial technology that enables human–computer interaction through natural language processing. In the context of the Bengali language, developing an accurate and efficient speech recognition system remains challenging due to its phonetic complexities and limited available resources. This research presents a comprehensive study on enhancing Bengali speech recognition by leveraging spectrogram analysis and deep learning techniques. We first investigate the fundamental aspects of spectrogram analysis to extract meaningful acoustic features from the Bengali speech signals. Next, we explore the efficacy of deep learning (DL) models using convolutional neural networks (CNNs) in capturing intricate linguistic patterns from spectrogram representations. Our research involves the construction of a medium-scale Bengali speech dataset, facilitating rigorous evaluation of the proposed methods. The experimental outcomes illustrate the effectiveness of our technique, showcasing significant advancements in Bengali speech recognition accuracy and performance. The findings presented in this paper hold promising implications for various applications, including voice-controlled systems, language translation, and transcription services in the Bengali-speaking community.