Fake speeches generated from text-to-speech (TTS) and voice-conversion (VC) technologies presents serious security threats, including impersonation, fraud, and misinformation. Fake speech detection is increasing important to tackle these threats. Traditional methods treat the task as a binary classification task, optimizing for Minimum Classification Error (MCE), but often failing to handle diverse fake speech effectively and lacking interpretability. In this work, we solve the task by maximizing the detection of the target class (i.e. fake speech class) with high precision, and considering all non-detected instances as the non-target class (i.e. genuine speech class). We propose to construct a novel multi-detector decision forest for fake speech detection. The method is composed of two stages, detecting stage and deciding stage. The detecting stage includes computing sub-task detections by multiple detectors trained with Adaptive Precision-Aimed Training procedure, which optimizes for Maximum Detection Precision (MDP). The deciding stage combines the sub-task detectors’ results through an interpretable way, which adopts several decision trees and OR logic calculus to form a decision forest. Our method achieve a superior performance with an Accuracy of 0.9820 on ASVspoof 2019 Logical Access (LA) dataset. Experiments also demonstrate that our method can effectively detect various types of unknown fake speech.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Constructing Multi-detector Decision Forest for Fake Speech Detection

  • Chang Feng,
  • Xiaolong Wu,
  • Yiyang Zhao,
  • Mingxing Xu,
  • Thomas Fang Zheng

摘要

Fake speeches generated from text-to-speech (TTS) and voice-conversion (VC) technologies presents serious security threats, including impersonation, fraud, and misinformation. Fake speech detection is increasing important to tackle these threats. Traditional methods treat the task as a binary classification task, optimizing for Minimum Classification Error (MCE), but often failing to handle diverse fake speech effectively and lacking interpretability. In this work, we solve the task by maximizing the detection of the target class (i.e. fake speech class) with high precision, and considering all non-detected instances as the non-target class (i.e. genuine speech class). We propose to construct a novel multi-detector decision forest for fake speech detection. The method is composed of two stages, detecting stage and deciding stage. The detecting stage includes computing sub-task detections by multiple detectors trained with Adaptive Precision-Aimed Training procedure, which optimizes for Maximum Detection Precision (MDP). The deciding stage combines the sub-task detectors’ results through an interpretable way, which adopts several decision trees and OR logic calculus to form a decision forest. Our method achieve a superior performance with an Accuracy of 0.9820 on ASVspoof 2019 Logical Access (LA) dataset. Experiments also demonstrate that our method can effectively detect various types of unknown fake speech.