<p>Speaker verification systems deployed at the network edge face significant challenges in adverse acoustic environments. To address this, we propose a spatial-temporal graph attention network (ST-GAT) framework for AI-Flow-compliant speaker verification using ad-hoc microphone arrays, where distributed nodes function as intelligent edge endpoints. Specifically, it includes a distributed feature aggregation block and an adaptive channel selection block, both built on graph structures that enable efficient cooperation among spatially distributed microphone nodes. The feature aggregation block fuses speaker features among different time and channels by a ST-GAT. The graph-based channel selection block chooses a good subset of microphones that may contribute apparently to the performance improvement of the system. The proposed method is flexible in incorporating various kinds of graphs and prior knowledge. We compared the proposed method with six representative methods in both real-world and simulated environments. Experimental results show that the proposed method achieves a relative equal error rate (EER) reduction of 15.39% lower than the strongest referenced method in the simulated datasets, and 17.70% lower than the latter in the real datasets. Moreover, its performance is robust across different signal-to-noise ratios and reverberation time.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Edge-collaborative multi-channel speaker verification via spatial-temporal graph with ad-hoc microphone arrays

  • Yijiang Chen,
  • Chengdong Liang,
  • Sizhou Chen,
  • Linfeng Feng,
  • Boyu Zhu,
  • Chi Zhang,
  • Xiao-Lei Zhang

摘要

Speaker verification systems deployed at the network edge face significant challenges in adverse acoustic environments. To address this, we propose a spatial-temporal graph attention network (ST-GAT) framework for AI-Flow-compliant speaker verification using ad-hoc microphone arrays, where distributed nodes function as intelligent edge endpoints. Specifically, it includes a distributed feature aggregation block and an adaptive channel selection block, both built on graph structures that enable efficient cooperation among spatially distributed microphone nodes. The feature aggregation block fuses speaker features among different time and channels by a ST-GAT. The graph-based channel selection block chooses a good subset of microphones that may contribute apparently to the performance improvement of the system. The proposed method is flexible in incorporating various kinds of graphs and prior knowledge. We compared the proposed method with six representative methods in both real-world and simulated environments. Experimental results show that the proposed method achieves a relative equal error rate (EER) reduction of 15.39% lower than the strongest referenced method in the simulated datasets, and 17.70% lower than the latter in the real datasets. Moreover, its performance is robust across different signal-to-noise ratios and reverberation time.