A Transformer-Based Depression Detection Network Leveraging Speech Emotional Expression Cues
摘要
In recent years, significant progress has been made in automated depression detection methods using speech and text data combined with deep learning. However, few studies have explored the connection between depression and speech emotions. To address this issue, this paper proposes a novel Transformer-based network leveraging Speech Emotion Information (SEI) for depression detection. The proposed network consists of a primary network used for depression classification and an auxiliary network is employed for speech emotion classification. In the primary network, the pre-trained Hubert and RoBERTa are used to obtain the short-term acoustic and textual features, respectively. And then the long-term audio and text features are aggregated from the shot-term features by using Transformer-based approaches with average pooling. The SEI extracted from the auxiliary network serves as supplementary auxiliary features aimed at augmenting the precision of depression recognition. Based on the proposed method, our best experimental results achieved an accuracy of 76.10% at the subject level. The experimental results demonstrate that incorporating speech emotions in depression detection improves diagnostic accuracy, offering a new perspective for research in this area.