Empowering Bangla Speech Recognition System Through Spectrogram Analysis and Deep Learning Approach
摘要
Speech recognition is a crucial technology that enables human–computer interaction through natural language processing. In the context of the Bengali language, developing an accurate and efficient speech recognition system remains challenging due to its phonetic complexities and limited available resources. This research presents a comprehensive study on enhancing Bengali speech recognition by leveraging spectrogram analysis and deep learning techniques. We first investigate the fundamental aspects of spectrogram analysis to extract meaningful acoustic features from the Bengali speech signals. Next, we explore the efficacy of deep learning (DL) models using convolutional neural networks (CNNs) in capturing intricate linguistic patterns from spectrogram representations. Our research involves the construction of a medium-scale Bengali speech dataset, facilitating rigorous evaluation of the proposed methods. The experimental outcomes illustrate the effectiveness of our technique, showcasing significant advancements in Bengali speech recognition accuracy and performance. The findings presented in this paper hold promising implications for various applications, including voice-controlled systems, language translation, and transcription services in the Bengali-speaking community.