Graph Attention-assisted Encoding for Time-domain Speech Separation
摘要
Deep learning-based time-domain speech separation methods have garnered much attention due to their remarkable effects. In addition to the key role of the separator in the time-domain separation model, the choice of encoder and decoder also significantly impacts separation performance. An overly large sliding window in the encoder can blur local speech features, diminishing separation performance. Conversely, a smaller window can enhance separation performance, but results in longer front-end representations, which increases computational costs and poses challenges in modeling temporal dependencies in long speech sequences. Here, we propose a speech separation network based on graph attention network-assisted encoding and further expand the exploration of the encoder’s role and responsibilities. It utilizes the graph network to learn the structural relationships between speech nodes, thereby compensating for the coarse-grained information under large windows. This ensures that the separation network effectively models short and low-resolution encoded sequences, improving computational efficiency and reducing training costs. Specifically, after constructing a graph representation for every latent feature, the graph attention network re-aggregates the features by assigning different learning weights to adjacent node features. The transformed new features can better guide the separation and reconstruction of speech. In extensive experiments on benchmark datasets, the proposed method outperforms other methods for time-domain speech encoders, offering faster inference speed and significantly improving the quality of separated speech.