On the Influence of CNN-Based Feature Learning Modules in Neural Speaker Verification Framework
摘要
The advent of Convolutional Neural Networks (CNNs) has revolutionized various fields of machine learning, including automatic speaker verification (ASV). In this paper, we explore the influence of CNN-based feature learning modules on the performance and robustness of neural ASV systems. To this end, we employ a neural speaker embedding framework consisting of CNN-based features extraction module, which is connected in cascade with a frame-level network, consisting of a Time Delay Neural Network (TDNN)-Long Short Term Memory (LSTM) hybrid network and a fully TDNN network in a cascade arrangement. The CNN-based feature extraction modules considered are: (i) 1-D CNN, (ii) vanilla 2D-CNN, and (iii) Selective Kernel Attention (SKA) integrated 2D-CNN. In order to aggregate speaker information within an utterance-level context, we use Multi-Level Attentive Statistics Pooling, which captures local statistics as context and leverages the complementarity of different network modules. Experimental results on the Voxceleb dataset show that the frequency- and channel-wise SKA integrated 2D-CNN-based feature extraction module significantly outperforms the other considered feature extraction modules.