A New Hybrid VAD Model Using Machine Learning System and GFCC Acoustic Features in Challenging Noisy Environment
摘要
Recently, machine learning (ML)-based voice activity detection (VAD) has achieved satisfactory performance in stationary noise, while the presence of transients in audio signals still poses an interesting challenge in the VAD domain. To deal with this problem, we propose a new hybrid VAD model that combines supervised ML classifiers and leverages gammatone frequency cepstral coefficients (GFCC) as relevant parameters to distinguish between speech and non-speech frames corrupted by stationary and non-stationary noises. The proposed approach consists of two sequential stages. In the first stage, the partial decisions provided by two ML classifiers are combined through a fusion rule to enhance the VAD model accuracy in stationary noise situations. In particular, this involves associating Pseudo-Quadratic Discriminant Analysis and Support Vector Machine by leveraging the advantages of each classifier. The second stage consists of the modified bidirectional-long short-term memory (Bi-LSTM) network, which is directly fed with the output of the first stage along with the GFCC features. Subsequently, the smoothing procedure is performed to make the final VAD decision. The proposed model has been evaluated through LibriSpeech and TIMIT databases contaminated by both stationary and non-stationary noise at different signal-to-noise ratios. Through experimental evaluation, the hybrid VAD model not only improves performance in stationary noise but also acts more robustly in challenging noisy environments compared to using each classifier separately. The proposed method achieves 94.21% and 96.09% accuracy in stationary and non-stationary noise conditions, respectively, slightly exceeding recent models for the same noisy situations.