Speech has been adopted as a bioinformatic for detecting depression. Current speech-based depression detection (SDD) methods rely on raw signal and Mel-scale features as input. However, utilizing raw signal as input leads to high model complexity, while using Mel-scale features reduces model performance due to insufficient domain-specific adaptation. To solve these issues, we present a depression speech analysis (DSPA) network, which contains a task-oriented learnable frequency-domain filterbanks (LFB) module for optimizing spectral feature generation via end-to-end tuning of filter parameters, and a spectro-temporal representation extraction (STE) module for identifying depression representations in the LFB learned features, while guiding filter parameters optimization. Furthermore, due to LFB exhibiting sensitivity to parameter initialization, a speech representation disentanglement strategy is designed to guide filters to focus on the emotional representations, where pre-trained parameters are used for initializing LFB. Our method yields F1-scores of 0.792, 0.927 and 0.702 on the DAIC-woz, CMDC and EATD-corpus datasets, respectively, with an average improvement of 8.9%, 1.2% and 8.7% over compared state-of-the-art methods. The results show that DSPA is effective in extracting depression-related features.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Optimizing Learnable Frequency-Domain Filterbanks for Depression Detection via Speech Representation Disentanglement

  • Wenju Yang,
  • LeYang Li,
  • Yong Hao,
  • Peng Cao,
  • Osmar R. Zaiane

摘要

Speech has been adopted as a bioinformatic for detecting depression. Current speech-based depression detection (SDD) methods rely on raw signal and Mel-scale features as input. However, utilizing raw signal as input leads to high model complexity, while using Mel-scale features reduces model performance due to insufficient domain-specific adaptation. To solve these issues, we present a depression speech analysis (DSPA) network, which contains a task-oriented learnable frequency-domain filterbanks (LFB) module for optimizing spectral feature generation via end-to-end tuning of filter parameters, and a spectro-temporal representation extraction (STE) module for identifying depression representations in the LFB learned features, while guiding filter parameters optimization. Furthermore, due to LFB exhibiting sensitivity to parameter initialization, a speech representation disentanglement strategy is designed to guide filters to focus on the emotional representations, where pre-trained parameters are used for initializing LFB. Our method yields F1-scores of 0.792, 0.927 and 0.702 on the DAIC-woz, CMDC and EATD-corpus datasets, respectively, with an average improvement of 8.9%, 1.2% and 8.7% over compared state-of-the-art methods. The results show that DSPA is effective in extracting depression-related features.