Edge-collaborative multi-channel speaker verification via spatial-temporal graph with ad-hoc microphone arrays
摘要
Speaker verification systems deployed at the network edge face significant challenges in adverse acoustic environments. To address this, we propose a spatial-temporal graph attention network (ST-GAT) framework for AI-Flow-compliant speaker verification using ad-hoc microphone arrays, where distributed nodes function as intelligent edge endpoints. Specifically, it includes a distributed feature aggregation block and an adaptive channel selection block, both built on graph structures that enable efficient cooperation among spatially distributed microphone nodes. The feature aggregation block fuses speaker features among different time and channels by a ST-GAT. The graph-based channel selection block chooses a good subset of microphones that may contribute apparently to the performance improvement of the system. The proposed method is flexible in incorporating various kinds of graphs and prior knowledge. We compared the proposed method with six representative methods in both real-world and simulated environments. Experimental results show that the proposed method achieves a relative equal error rate (EER) reduction of 15.39% lower than the strongest referenced method in the simulated datasets, and 17.70% lower than the latter in the real datasets. Moreover, its performance is robust across different signal-to-noise ratios and reverberation time.