For speaker recognition, in addition to recognition accuracy, large-scale speaker recognition faces another challenge: fast search of speaker databases. In this paper, we propose a voiceprint recognition system based on a deep neural network recognition method and locality-sensitive hashing (high-dimensional fast nearest neighbor search algorithm). The embedding vector representing the speaker’s characteristics is extracted through the deep neural network to form a registered voiceprint library, and then the voiceprint library is hash-encoded using local sensitive hashing, so that the high-dimensional embedding vector is mapped into a one-dimensional hash code. At the same time, these hash codes retain the similarity features between the original voiceprint embedding vectors. Using this method for large-scale speaker identification can significantly reduce retrieval time. We evaluate our method on Aishell-2, a real-world dataset containing approximately 2,000 speakers. The results show that the algorithm proposed in this article is approximately 300 times faster than conventional retrieval methods, while ensuring the recognition accuracy of the voiceprint recognition system.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Speaker Recognition Based on Locality Sensitive Hashing

  • Yifan Wu,
  • Erhua Zhang,
  • Chunxia Hou,
  • Zhenmin Tang

摘要

For speaker recognition, in addition to recognition accuracy, large-scale speaker recognition faces another challenge: fast search of speaker databases. In this paper, we propose a voiceprint recognition system based on a deep neural network recognition method and locality-sensitive hashing (high-dimensional fast nearest neighbor search algorithm). The embedding vector representing the speaker’s characteristics is extracted through the deep neural network to form a registered voiceprint library, and then the voiceprint library is hash-encoded using local sensitive hashing, so that the high-dimensional embedding vector is mapped into a one-dimensional hash code. At the same time, these hash codes retain the similarity features between the original voiceprint embedding vectors. Using this method for large-scale speaker identification can significantly reduce retrieval time. We evaluate our method on Aishell-2, a real-world dataset containing approximately 2,000 speakers. The results show that the algorithm proposed in this article is approximately 300 times faster than conventional retrieval methods, while ensuring the recognition accuracy of the voiceprint recognition system.