Multi-path Processing Structure for Multi-scale Feature Fusion in Speech Separation Transformer
摘要
Recent advancements in time-domain audio separation networks (TasNets) have markedly propelled the field of speech separation. Unlike conventional time-frequency domain methodologies, TasNets directly model the amalgamated speech signals in the time-domain, employing a convolutional encoder-decoder architecture to perform separation on the output of the encoder. However, the original dual-path framework is characterized by a fixed feature dimension and a constant segment size across all recurrent neural network layers, thereby limiting its ability to generate high-resolution features. In this study, we present a novel approach termed Multi-Scale Feature Fusion Transformer Network (MSFFT-Net). The MSFFT-Net incorporates multiple dual-path processing paths in the separation stage, with each path dedicates to perform feature modeling at different scales. Coarse-grain and fine-grain features are obtained in parallel from different processing paths. Additionally, the features from one dual-path processing path can be exchanged and shared with other distinct processing path, ultimately yielding high feature resolution across layers, thereby enhancing mask estimation accuracy. Experimental results on several datasets demonstrate the superiority of the MSFFT-Net over SOTA baselines in the single channel speech separation task.