Hybrid Convolutional Transposed Attention Mechanism Based on HOII Former Network for Speech Enhancement
摘要
Recent deep learning based speech enhancement models highlight the effectiveness of attention mechanism in state of the art approaches. Convolutional Neural Networks with fixed kernel sizes capture temporal modules. Transformer techniques surpass conventional neural networks, such as RNNs, and CNNs, in handling long-term dependencies. To address current limitations, the proposed Sub-convolutional Squeezed Hybrid Gated Mechanism-Fusion-Dual-path HOII Former Network (SCGM-DPH-Net) integrates several advanced components. Its sub convolutional encoder-decoder with variable kernel sizes extracts multi-scale information from noisy speech. STCM blocs model spectral and temporal sequences, and the hybrid convolutional transposed attention mechanism improves local–global feature correlations. The fusion function enhances multi-scale feature extraction. The Dual-Path HOII Former (DPH) models long-range dependencies through high-order time–frequency interactions. Experiments a how the proposed model outperforms many state-of-the-art approaches achieving higher average PESQ, SDR and STOI scores than the noisy speech model in various environments.