In recent years, significant progress has been made in automated depression detection methods using speech and text data combined with deep learning. However, few studies have explored the connection between depression and speech emotions. To address this issue, this paper proposes a novel Transformer-based network leveraging Speech Emotion Information (SEI) for depression detection. The proposed network consists of a primary network used for depression classification and an auxiliary network is employed for speech emotion classification. In the primary network, the pre-trained Hubert and RoBERTa are used to obtain the short-term acoustic and textual features, respectively. And then the long-term audio and text features are aggregated from the shot-term features by using Transformer-based approaches with average pooling. The SEI extracted from the auxiliary network serves as supplementary auxiliary features aimed at augmenting the precision of depression recognition. Based on the proposed method, our best experimental results achieved an accuracy of 76.10% at the subject level. The experimental results demonstrate that incorporating speech emotions in depression detection improves diagnostic accuracy, offering a new perspective for research in this area.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A Transformer-Based Depression Detection Network Leveraging Speech Emotional Expression Cues

  • Changqing Xu,
  • Xinyi Wu,
  • Nan Li,
  • Xin Wang,
  • Xu Feng,
  • Rongfeng Su,
  • Nan Yan,
  • Lan Wang

摘要

In recent years, significant progress has been made in automated depression detection methods using speech and text data combined with deep learning. However, few studies have explored the connection between depression and speech emotions. To address this issue, this paper proposes a novel Transformer-based network leveraging Speech Emotion Information (SEI) for depression detection. The proposed network consists of a primary network used for depression classification and an auxiliary network is employed for speech emotion classification. In the primary network, the pre-trained Hubert and RoBERTa are used to obtain the short-term acoustic and textual features, respectively. And then the long-term audio and text features are aggregated from the shot-term features by using Transformer-based approaches with average pooling. The SEI extracted from the auxiliary network serves as supplementary auxiliary features aimed at augmenting the precision of depression recognition. Based on the proposed method, our best experimental results achieved an accuracy of 76.10% at the subject level. The experimental results demonstrate that incorporating speech emotions in depression detection improves diagnostic accuracy, offering a new perspective for research in this area.