Dimensionality Reduction in Multicepstral Features for Voice Spoofing Detection: Case Studies with Singular Value Decomposition, Genetic Algorithm, and Auto-Encoder
摘要
Recognizing people by their voice is not only interesting and challenging on its own, but also a very necessary ability in a variety of contexts, e.g. for unlocking electronic devices, for authentication in banking transactions and much more. However, current automatic voice verification technologies are vulnerable to presentation attacks due to the increasing realm of audio falsifications – also known as spoofing – that rely on artificial intelligence technologies, for example. Due to the direct connection of speech recognition with security systems and privacy concerns, the development of measures against spoofing is not only crucial but also urgently needed. In this work, we propose an experimental approach that focuses on dimensionality reduction techniques together with a classification model to detect spoofing in voice biometric systems. Our method uses a multicepstral feature extraction framework to distinguish between real and synthetic speech signals. To validate the proposed method, tests were performed using the ASVSpoof 2017 v2.0 database. Dimensionality reduction techniques such as singular value decomposition, genetic algorithms and auto-encoder were applied before implementing the support vector machine model. The model is then evaluated using the Equal Error Rate metric. For comparison, we investigated the voice liveness detection problem using machine learning models with and without dimensionality reduction strategies, observing an improvement of up to \(7.11\%\) in the model error metric.