This study proposes a method to detect audio deepfakes by leveraging the inherent difficulty of replicating human emotions using deep learning models. We introduce a detection framework that utilizes Valence-Arousal-Dominance (VAD), which estimates emotional states in speech and extracts auxiliary features. These features, combined with latent embeddings from an efficient speech representation model, significantly improve detection accuracy. Our experimental evaluation demonstrates that the proposed approach enhances state-of-the-art methods like AASIST across multiple languages and environmental settings. This research suggests that incorporating emotional cues could be a promising approach to tackling the evolving challenge of audio deepfakes.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Detecting Audio Deepfakes Through Emotional Fingerprinting

  • Long Nguyen-Vu,
  • Yowon Lee,
  • Thien-Phuc Doan,
  • Kihun Hong,
  • Souhwan Jung

摘要

This study proposes a method to detect audio deepfakes by leveraging the inherent difficulty of replicating human emotions using deep learning models. We introduce a detection framework that utilizes Valence-Arousal-Dominance (VAD), which estimates emotional states in speech and extracts auxiliary features. These features, combined with latent embeddings from an efficient speech representation model, significantly improve detection accuracy. Our experimental evaluation demonstrates that the proposed approach enhances state-of-the-art methods like AASIST across multiple languages and environmental settings. This research suggests that incorporating emotional cues could be a promising approach to tackling the evolving challenge of audio deepfakes.