Bangla (Bengali), one of the world’s most spoken languages, presents significant challenges for Automatic Speech Recognition (ASR) due to its diverse yet underrepresented dialects. Existing speech technology often underperforms on these variations, and dedicated solutions for Bangla dialect identification (DID) remain underdeveloped. This paper presents a novel hybrid deep learning approach tailored for Bangla DID, using an expanded dataset covering major dialectal regions of Bangladesh. We utilize this approach to develop a robust DID model and offer insights into effective modeling strategies for Bangla’s diverse dialects. Our main contributions are as follows: (i) We curated and utilized an expanded dataset covering six major Bangladeshi dialectal regions by merging existing resources with newly collected data, addressing critical data scarcity and imbalance issues previously hindering Bangla DID research. (ii) We designed the innovative hybrid GLDNN architecture, sequencing Gated Recurrent Unit (GRU) and Long Short-Term Memory (LSTM) layers, to effectively capture both the short-term phonetic details and longer-term prosodic patterns inherent in Bangla dialectal speech variations. (iii) We demonstrate through systematic evaluation that both audio segment length and oversampling are critical factors for Bangla DID performance; shorter segments (3–5 s) proved most effective, significantly impacting accuracy and highlighting the interplay between temporal context resolution and data balance in this low-resource setting. (iv) Through extensive experiments, we validate the efficacy of the GLDNN approach, showing it significantly outperforms classical baselines and standalone recurrent architectures, achieving up to 98% accuracy on benchmark configurations and demonstrating robust, balanced performance across diverse dialects.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A Novel Hybrid GLDNN Architecture for Bangla Dialect Identification

  • Mohammad Arafath Uddin Shariff

摘要

Bangla (Bengali), one of the world’s most spoken languages, presents significant challenges for Automatic Speech Recognition (ASR) due to its diverse yet underrepresented dialects. Existing speech technology often underperforms on these variations, and dedicated solutions for Bangla dialect identification (DID) remain underdeveloped. This paper presents a novel hybrid deep learning approach tailored for Bangla DID, using an expanded dataset covering major dialectal regions of Bangladesh. We utilize this approach to develop a robust DID model and offer insights into effective modeling strategies for Bangla’s diverse dialects. Our main contributions are as follows: (i) We curated and utilized an expanded dataset covering six major Bangladeshi dialectal regions by merging existing resources with newly collected data, addressing critical data scarcity and imbalance issues previously hindering Bangla DID research. (ii) We designed the innovative hybrid GLDNN architecture, sequencing Gated Recurrent Unit (GRU) and Long Short-Term Memory (LSTM) layers, to effectively capture both the short-term phonetic details and longer-term prosodic patterns inherent in Bangla dialectal speech variations. (iii) We demonstrate through systematic evaluation that both audio segment length and oversampling are critical factors for Bangla DID performance; shorter segments (3–5 s) proved most effective, significantly impacting accuracy and highlighting the interplay between temporal context resolution and data balance in this low-resource setting. (iv) Through extensive experiments, we validate the efficacy of the GLDNN approach, showing it significantly outperforms classical baselines and standalone recurrent architectures, achieving up to 98% accuracy on benchmark configurations and demonstrating robust, balanced performance across diverse dialects.