With the increasing concern about network security and model interpretability, fake audio detection is attracting significant attention. Although prior studies primarily concentrate on distinguishing authentic from fabricated audio, a vital yet underexplored dimension is the pinpointing within such detections, particularly in identifying the impersonated individual from counterfeit audio. In this paper, we propose a model capable of simultaneously detecting fake audio and identifying the speaker by extracting ID-specific features. Furthermore, we produce the Fake In-the-Wild Audio (FIA) dataset by expanding the “In-the-Wild” audio dataset. We adopt the advanced Text-to-Speech generation model, MetaVoice-1B, to generate fake audios based on the “In-the-Wild” audio dataset. We conducted a detailed analysis of MetaVoice-1B, focusing on its capabilities in generating realistic deepfake audio. Additionally, we identified its disadvantages and suggest potential future improvements. Experiments on the FIA dataset demonstrate the excellent performance of the proposed model, achieving a remarkable F1 score of 0.99 on the test set. Moreover, the model’s ability to output additional relevant information enhances overall cybersecurity by providing deeper insights into potentially fake content.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Who is Being Impersonated? Deepfake Audio Detection and Impersonated Identification via Extraction of Id-Specific Features

  • Tianchen Guo,
  • Heming Du,
  • Huan Huo,
  • Bo Liu,
  • Xin Yu

摘要

With the increasing concern about network security and model interpretability, fake audio detection is attracting significant attention. Although prior studies primarily concentrate on distinguishing authentic from fabricated audio, a vital yet underexplored dimension is the pinpointing within such detections, particularly in identifying the impersonated individual from counterfeit audio. In this paper, we propose a model capable of simultaneously detecting fake audio and identifying the speaker by extracting ID-specific features. Furthermore, we produce the Fake In-the-Wild Audio (FIA) dataset by expanding the “In-the-Wild” audio dataset. We adopt the advanced Text-to-Speech generation model, MetaVoice-1B, to generate fake audios based on the “In-the-Wild” audio dataset. We conducted a detailed analysis of MetaVoice-1B, focusing on its capabilities in generating realistic deepfake audio. Additionally, we identified its disadvantages and suggest potential future improvements. Experiments on the FIA dataset demonstrate the excellent performance of the proposed model, achieving a remarkable F1 score of 0.99 on the test set. Moreover, the model’s ability to output additional relevant information enhances overall cybersecurity by providing deeper insights into potentially fake content.